Abstract
dc:descriptionState-of-the-art machine learning models are extremely powerful and are finally breaking through into commercial products for computer vision and natural language processing. One common factor among these successful models is, they all require massive datasets for training. Following this trend, large-scale learning-based methods present a promising way forward for robotics research. This line of thinking naturally raises two questions: from where can we collect the appropriate data? and, how can it be leveraged to create effective robotic systems? Fortunately, a vast amount of data already exists, showcasing the complexity of real-world environments and interactions that robots need to understand, in the form of videos. However, these video sources of data cannot be directly used with the techniques traditionally applied for robot learning. Videos may lack explicit action or goal labels, often depict suboptimal trajectories, and present a significant embodiment gap, both visually and in dynamics. These challenges underscore the need for new robot learning methods that can overcome these obstacles. In this work, we present our efforts to realize the goal of robot learning at scale using in-the-wild videos, developing methods to address each of the challenges that limit robot learning from videos. We introduce techniques for inferring actions and goals in unlabelled video data, learning optimal behavior from sub-optimal data, and tackling the embodiment gap by leveraging factored representations. Overall, this dissertation lays the foundations for how video data can be leveraged for robot learning at scale. We hope this work can serve as a step towards general robotic agents that can make significant positive impacts in the world.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chang, Matthew
- Contributors dc:contributor
-
- Gupta, Saurabh
- Forsyth, David
- Lazebnik, Svetlana
- Chaplot, Devendra
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Matthew Chang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/124136