{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127269"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127269","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards adaptive voice-controlled robots","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_has_math":false,"creators":["Chang, Peixin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Driggs-Campbell, Katherine","Hockenmaier, Julia","Chowdhary, Girish","Gupta, Saurabh"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-05","date_published":"2024-12-05","updated_at":"2026-07-22T22:25:03Z","subjects":["Robotics","Representation Learning","Continual Learning","Reinforcement Learning","Speech Recognition"],"languages":["en","eng"],"rights":["Copyright 2024 Peixin Chang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127269","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Driggs-Campbell, Katherine","Hockenmaier, Julia","Chowdhary, Girish","Gupta, Saurabh"]},{"key":"dc:creator","label":"Author","values":["Chang, Peixin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-12-05","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Robotics","Representation Learning","Continual Learning","Reinforcement Learning","Speech Recognition"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Peixin Chang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127269"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Peixin Chang, accepted the attached license on 2024-12-04 at 22:34.","The student, Peixin Chang, submitted this Dissertation for approval on 2024-12-04 at 22:55.","This Dissertation was approved for publication on 2024-12-05 at 10:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21498 on 2025-03-28 at 14:28:03","Voice-controlled robots offer a natural interface for non-experts to communicate with robots. Previous methods either rely on modular pipelines, which suffer from cascading errors and poor integration between components, or on end-to-end models that struggle with robustness and generalization, especially when applied to new tasks or environments. Both approaches often demand significant domain expertise and extensive manual tuning after deployment if failures or suboptimal behaviors occur, making them difficult for non-experts to update or adapt the system. These drawbacks undermine seamless human-robot collaboration and limit the adoption of voice-controlled robots in daily life. In this thesis, we address the challenge of enabling voice-controlled robots to adapt and improve after deployment with minimal supervision and assumptions, a challenge frequently overlooked in current research and development. To this end, we propose a novel two-stage pipeline. The first stage involves developing a Visual-Audio Representation (VAR), which unifies speech recognition, natural language understanding, and grounding modules, allowing the robot to ground multimodal inputs. The second stage employs a reinforcement learning policy that uses the learned embeddings and rewards from the VAR, enabling the robot to improve after the deployment. We also introduce Dif-VAR, a data-efficient version of the VAR that allows for intuitive fine-tuning by non-experts with significantly reduced labeling requirements. Our system has been rigorously evaluated using state-of-the-art sound datasets in both simulated and real-world environments, showing robust performance improvements in various navigation and manipulation tasks. The proposed approach allows for continual self-improvement of the robot after the deployment, with minimal data and human intervention, making it a scalable and adaptive solution for voice-controlled robots in everyday scenarios."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards adaptive voice-controlled robots"]}]}],"canonical_facts":{"dc:contributor":["Driggs-Campbell, Katherine","Hockenmaier, Julia","Chowdhary, Girish","Gupta, Saurabh"],"dc:creator":["Chang, Peixin"],"dc:date":["2024-12-05","2024-12"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Peixin Chang, accepted the attached license on 2024-12-04 at 22:34.","The student, Peixin Chang, submitted this Dissertation for approval on 2024-12-04 at 22:55.","This Dissertation was approved for publication on 2024-12-05 at 10:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21498 on 2025-03-28 at 14:28:03","Voice-controlled robots offer a natural interface for non-experts to communicate with robots. Previous methods either rely on modular pipelines, which suffer from cascading errors and poor integration between components, or on end-to-end models that struggle with robustness and generalization, especially when applied to new tasks or environments. Both approaches often demand significant domain expertise and extensive manual tuning after deployment if failures or suboptimal behaviors occur, making them difficult for non-experts to update or adapt the system. These drawbacks undermine seamless human-robot collaboration and limit the adoption of voice-controlled robots in daily life. In this thesis, we address the challenge of enabling voice-controlled robots to adapt and improve after deployment with minimal supervision and assumptions, a challenge frequently overlooked in current research and development. To this end, we propose a novel two-stage pipeline. The first stage involves developing a Visual-Audio Representation (VAR), which unifies speech recognition, natural language understanding, and grounding modules, allowing the robot to ground multimodal inputs. The second stage employs a reinforcement learning policy that uses the learned embeddings and rewards from the VAR, enabling the robot to improve after the deployment. We also introduce Dif-VAR, a data-efficient version of the VAR that allows for intuitive fine-tuning by non-experts with significantly reduced labeling requirements. Our system has been rigorously evaluated using state-of-the-art sound datasets in both simulated and real-world environments, showing robust performance improvements in various navigation and manipulation tasks. The proposed approach allows for continual self-improvement of the robot after the deployment, with minimal data and human intervention, making it a scalable and adaptive solution for voice-controlled robots in everyday scenarios."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127269"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Peixin Chang"],"dc:subject":["Robotics","Representation Learning","Continual Learning","Reinforcement Learning","Speech Recognition"],"dc:title":["Towards adaptive voice-controlled robots"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:03Z"}