Abstract
dc:descriptionVoice-controlled robots offer a natural interface for non-experts to communicate with robots. Previous methods either rely on modular pipelines, which suffer from cascading errors and poor integration between components, or on end-to-end models that struggle with robustness and generalization, especially when applied to new tasks or environments. Both approaches often demand significant domain expertise and extensive manual tuning after deployment if failures or suboptimal behaviors occur, making them difficult for non-experts to update or adapt the system. These drawbacks undermine seamless human-robot collaboration and limit the adoption of voice-controlled robots in daily life. In this thesis, we address the challenge of enabling voice-controlled robots to adapt and improve after deployment with minimal supervision and assumptions, a challenge frequently overlooked in current research and development. To this end, we propose a novel two-stage pipeline. The first stage involves developing a Visual-Audio Representation (VAR), which unifies speech recognition, natural language understanding, and grounding modules, allowing the robot to ground multimodal inputs. The second stage employs a reinforcement learning policy that uses the learned embeddings and rewards from the VAR, enabling the robot to improve after the deployment. We also introduce Dif-VAR, a data-efficient version of the VAR that allows for intuitive fine-tuning by non-experts with significantly reduced labeling requirements. Our system has been rigorously evaluated using state-of-the-art sound datasets in both simulated and real-world environments, showing robust performance improvements in various navigation and manipulation tasks. The proposed approach allows for continual self-improvement of the robot after the deployment, with minimal data and human intervention, making it a scalable and adaptive solution for voice-controlled robots in everyday scenarios.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chang, Peixin
- Contributors dc:contributor
-
- Driggs-Campbell, Katherine
- Hockenmaier, Julia
- Chowdhary, Girish
- Gupta, Saurabh
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Peixin Chang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/127269