{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/137800"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/137800","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots","abstract":"As robots transition from isolated industrial settings to working in close proximity to humans in dynamic environments, their ability to learn and adapt to human feedback and unseen circumstances becomes crucial. Imitation learning offers a promising paradigm for robots to learn complex tasks by mimicking human behavior. However, traditional imitation learning approaches face key challenges in integrating diverse feedback types, managing noisy and inconsistent inputs, and maintaining stability in learning. In this thesis, we develop imitation learning approaches that advance the capabilities of robots and enable them to efficiently learn from humans and adapt to unseen data in diverse environments. This research is structured around four key contributions. First, we consider the scenario where a human is readily available to provide high-quality feedback to the robot. We develop a learning algorithm to enable robots to learn from diverse sources of optimal human feedback: demonstrations, corrections, and preferences. Demonstrations provide high-level task overviews, corrections fine-tune specific motions, and preferences rank robot behaviors for task improvement. By incorporating these active and passive feedback sources under a unified reward learning framework we enable robots to infer task objectives more effectively and optimize their trajectories using constrained optimization techniques. Second, we explore scenarios where the human feedback is noisy or biased due to task complexity or physical constraints. We model the robot's learning rule as a dynamical system and apply Lyapunov stability analysis to derive conditions of convergence. Leveraging these conditions, we modify the robot's learning rule to expand the basins of attractions around the possible tasks (equilibrium points) in the environment. This approach enables the robot to infer the correct task representations from a wider range of human inputs making the learning robust to suboptimal feedback without destabilizing the robot behavior. Next, we consider imitation learning settings where a human is not available to provide additional feedback. In such scenarios, imitation learning algorithms are often prone to covariate shift when they encounter data not seen during training. To tackle this challenge, we develop Stable Behavior Cloning (Stable-BC), a stability-driven imitation learning algorithm. This algorithm ensures that robots maintain reliable performance by encouraging policy stability around demonstrated behaviors without the need for additional training data or complex reinforcement learning methods. Finally, we look at the problem of imitation learning from the users' perspective and aim to reduce the time and effort required to teach the robot. We propose L2D2, a sketching interface and imitation learning algorithm where humans can provide demonstrations by drawing the task. L2D2 leverages vision-language segmentation to autonomously vary object locations and generates synthetic images of the environment for the human to draw upon. By collecting a few physical demonstrations from the users, L2D2 then grounds these diverse 2D drawings in the real world. This approach reduces the time and effort required to teach the robots by enabling the users to rapidly provide a large set of diverse demonstrations. The findings from this research highlight the importance of adaptability and stability when robots and autonomous agents work around and interact with humans in diverse environments. This research contributes to the broader field of robot learning by offering scalable, adaptable, and user-friendly solutions for imitation learning and human-robot interaction, paving the way for more intuitive and robust robotic systems in human environments.","abstract_html":"As robots transition from isolated industrial settings to working in close proximity to humans in dynamic environments, their ability to learn and adapt to human feedback and unseen circumstances becomes crucial. Imitation learning offers a promising paradigm for robots to learn complex tasks by mimicking human behavior. However, traditional imitation learning approaches face key challenges in integrating diverse feedback types, managing noisy and inconsistent inputs, and maintaining stability in learning. In this thesis, we develop imitation learning approaches that advance the capabilities of robots and enable them to efficiently learn from humans and adapt to unseen data in diverse environments. This research is structured around four key contributions. First, we consider the scenario where a human is readily available to provide high-quality feedback to the robot. We develop a learning algorithm to enable robots to learn from diverse sources of optimal human feedback: demonstrations, corrections, and preferences. Demonstrations provide high-level task overviews, corrections fine-tune specific motions, and preferences rank robot behaviors for task improvement. By incorporating these active and passive feedback sources under a unified reward learning framework we enable robots to infer task objectives more effectively and optimize their trajectories using constrained optimization techniques. Second, we explore scenarios where the human feedback is noisy or biased due to task complexity or physical constraints. We model the robot&#x27;s learning rule as a dynamical system and apply Lyapunov stability analysis to derive conditions of convergence. Leveraging these conditions, we modify the robot&#x27;s learning rule to expand the basins of attractions around the possible tasks (equilibrium points) in the environment. This approach enables the robot to infer the correct task representations from a wider range of human inputs making the learning robust to suboptimal feedback without destabilizing the robot behavior. Next, we consider imitation learning settings where a human is not available to provide additional feedback. In such scenarios, imitation learning algorithms are often prone to covariate shift when they encounter data not seen during training. To tackle this challenge, we develop Stable Behavior Cloning (Stable-BC), a stability-driven imitation learning algorithm. This algorithm ensures that robots maintain reliable performance by encouraging policy stability around demonstrated behaviors without the need for additional training data or complex reinforcement learning methods. Finally, we look at the problem of imitation learning from the users&#x27; perspective and aim to reduce the time and effort required to teach the robot. We propose L2D2, a sketching interface and imitation learning algorithm where humans can provide demonstrations by drawing the task. L2D2 leverages vision-language segmentation to autonomously vary object locations and generates synthetic images of the environment for the human to draw upon. By collecting a few physical demonstrations from the users, L2D2 then grounds these diverse 2D drawings in the real world. This approach reduces the time and effort required to teach the robots by enabling the users to rapidly provide a large set of diverse demonstrations. The findings from this research highlight the importance of adaptability and stability when robots and autonomous agents work around and interact with humans in diverse environments. This research contributes to the broader field of robot learning by offering scalable, adaptable, and user-friendly solutions for imitation learning and human-robot interaction, paving the way for more intuitive and robust robotic systems in human environments.","abstract_has_math":false,"creators":["Mehta, Shaunak Abhijit"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Mechanical Engineering","degree_department":"Mechanical Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Losey, Dylan Patrick"],"committee_members":["Bajcsy, Andrea","Akbari Hamed, Kaveh","Bartlett, Michael David"],"year":2025,"date_issued":"2025-09-18","date_published":"2025-09-18","updated_at":"2026-07-22T22:20:25Z","subjects":["Imitation Learning","Human-Robot Interaction","Learning Dynamics"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44672"],"render_values":[{"text":"vt_gsexam:44672","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/137800","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Losey, Dylan Patrick"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Bajcsy, Andrea","Akbari Hamed, Kaveh","Bartlett, Michael David"]},{"key":"dc:contributor.department","label":"Department","values":["Mechanical Engineering"]},{"key":"dc:creator","label":"Author","values":["Mehta, Shaunak Abhijit"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-09-19T08:00:19Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-09-19T08:00:19Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-09-18"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechanical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Imitation Learning","Human-Robot Interaction","Learning Dynamics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44672"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/137800"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As robots transition from isolated industrial settings to working in close proximity to humans in dynamic environments, their ability to learn and adapt to human feedback and unseen circumstances becomes crucial. Imitation learning offers a promising paradigm for robots to learn complex tasks by mimicking human behavior. However, traditional imitation learning approaches face key challenges in integrating diverse feedback types, managing noisy and inconsistent inputs, and maintaining stability in learning. In this thesis, we develop imitation learning approaches that advance the capabilities of robots and enable them to efficiently learn from humans and adapt to unseen data in diverse environments. This research is structured around four key contributions. First, we consider the scenario where a human is readily available to provide high-quality feedback to the robot. We develop a learning algorithm to enable robots to learn from diverse sources of optimal human feedback: demonstrations, corrections, and preferences. Demonstrations provide high-level task overviews, corrections fine-tune specific motions, and preferences rank robot behaviors for task improvement. By incorporating these active and passive feedback sources under a unified reward learning framework we enable robots to infer task objectives more effectively and optimize their trajectories using constrained optimization techniques. Second, we explore scenarios where the human feedback is noisy or biased due to task complexity or physical constraints. We model the robot's learning rule as a dynamical system and apply Lyapunov stability analysis to derive conditions of convergence. Leveraging these conditions, we modify the robot's learning rule to expand the basins of attractions around the possible tasks (equilibrium points) in the environment. This approach enables the robot to infer the correct task representations from a wider range of human inputs making the learning robust to suboptimal feedback without destabilizing the robot behavior. Next, we consider imitation learning settings where a human is not available to provide additional feedback. In such scenarios, imitation learning algorithms are often prone to covariate shift when they encounter data not seen during training. To tackle this challenge, we develop Stable Behavior Cloning (Stable-BC), a stability-driven imitation learning algorithm. This algorithm ensures that robots maintain reliable performance by encouraging policy stability around demonstrated behaviors without the need for additional training data or complex reinforcement learning methods. Finally, we look at the problem of imitation learning from the users' perspective and aim to reduce the time and effort required to teach the robot. We propose L2D2, a sketching interface and imitation learning algorithm where humans can provide demonstrations by drawing the task. L2D2 leverages vision-language segmentation to autonomously vary object locations and generates synthetic images of the environment for the human to draw upon. By collecting a few physical demonstrations from the users, L2D2 then grounds these diverse 2D drawings in the real world. This approach reduces the time and effort required to teach the robots by enabling the users to rapidly provide a large set of diverse demonstrations. The findings from this research highlight the importance of adaptability and stability when robots and autonomous agents work around and interact with humans in diverse environments. This research contributes to the broader field of robot learning by offering scalable, adaptable, and user-friendly solutions for imitation learning and human-robot interaction, paving the way for more intuitive and robust robotic systems in human environments."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["As robots move beyond factory floors and start working alongside people in homes and workplaces, it becomes important for them to learn from human instructions and handle unexpected situations. One way robots can learn is by observing and copying human actions, similar to how children learn by watching adults. This approach is called imitation learning. While promising, it has challenges: robots need to handle different types of human instructions, deal with unclear or messy guidance, and keep their learning stable over time without making errors. This research focuses on creating smarter and more adaptable learning methods for robots to address these challenges. This research is organized around four main contributions. The first part of the research looks at situations where humans are available to give helpful guidance to robots. We develop a method that allows robots to learn from three types of human input: demonstrations (showing how to do a task), corrections (fixing mistakes), and preferences (choosing which robot action is better). By combining these types of input, robots can better understand what humans want and adjust their actions to complete tasks successfully. In the second part, we explore situations where human instructions are not perfect, perhaps due to the complexity of the task or physical limitations that affect their inputs. Sometimes, humans make mistakes or provide unclear feedback. To address this, we treat the robot's learning process as a system that needs to stay balanced and stable, even when the guidance is noisy. We develop a method that helps the robot understand the correct task from a wider range of human inputs without misinterpreting the inputs or becoming unstable. Next, we consider cases where a human isn't available to provide additional guidance during the robot's learning process. In such cases, robots often struggle when they face situations they haven't seen before. This can lead to the robot taking incorrect actions resulting in errors that get worse over time. To solve this problem, we create an approach called Stable Behavior Cloning, which helps robots stay consistent and perform reliably without needing additional training or complex trial-and-error learning. Finally, we think about the problem from the human user's point of view and try to make it easier and faster to teach robots. We introduce L2D2, a tool where people can teach robots by simply drawing what they want the robot to do. L2D2 uses advanced computer vision and language techniques to create different versions of the task environment automatically, so people can quickly provide many different examples. This makes teaching robots faster and less work for the user. This research shows that adaptability and stability are essential for robots to work safely and effectively alongside humans in diverse, real-world environments. The findings of this research help us move closer to building robots that are easier to use, safer, and more helpful in everyday life."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Losey, Dylan Patrick"],"dc:contributor.committeemember":["Bajcsy, Andrea","Akbari Hamed, Kaveh","Bartlett, Michael David"],"dc:contributor.department":["Mechanical Engineering"],"dc:creator":["Mehta, Shaunak Abhijit"],"dc:date.accessioned":["2025-09-19T08:00:19Z"],"dc:date.available":["2025-09-19T08:00:19Z"],"dc:date.issued":["2025-09-18"],"dc:description.abstract":["As robots transition from isolated industrial settings to working in close proximity to humans in dynamic environments, their ability to learn and adapt to human feedback and unseen circumstances becomes crucial. Imitation learning offers a promising paradigm for robots to learn complex tasks by mimicking human behavior. However, traditional imitation learning approaches face key challenges in integrating diverse feedback types, managing noisy and inconsistent inputs, and maintaining stability in learning. In this thesis, we develop imitation learning approaches that advance the capabilities of robots and enable them to efficiently learn from humans and adapt to unseen data in diverse environments. This research is structured around four key contributions. First, we consider the scenario where a human is readily available to provide high-quality feedback to the robot. We develop a learning algorithm to enable robots to learn from diverse sources of optimal human feedback: demonstrations, corrections, and preferences. Demonstrations provide high-level task overviews, corrections fine-tune specific motions, and preferences rank robot behaviors for task improvement. By incorporating these active and passive feedback sources under a unified reward learning framework we enable robots to infer task objectives more effectively and optimize their trajectories using constrained optimization techniques. Second, we explore scenarios where the human feedback is noisy or biased due to task complexity or physical constraints. We model the robot's learning rule as a dynamical system and apply Lyapunov stability analysis to derive conditions of convergence. Leveraging these conditions, we modify the robot's learning rule to expand the basins of attractions around the possible tasks (equilibrium points) in the environment. This approach enables the robot to infer the correct task representations from a wider range of human inputs making the learning robust to suboptimal feedback without destabilizing the robot behavior. Next, we consider imitation learning settings where a human is not available to provide additional feedback. In such scenarios, imitation learning algorithms are often prone to covariate shift when they encounter data not seen during training. To tackle this challenge, we develop Stable Behavior Cloning (Stable-BC), a stability-driven imitation learning algorithm. This algorithm ensures that robots maintain reliable performance by encouraging policy stability around demonstrated behaviors without the need for additional training data or complex reinforcement learning methods. Finally, we look at the problem of imitation learning from the users' perspective and aim to reduce the time and effort required to teach the robot. We propose L2D2, a sketching interface and imitation learning algorithm where humans can provide demonstrations by drawing the task. L2D2 leverages vision-language segmentation to autonomously vary object locations and generates synthetic images of the environment for the human to draw upon. By collecting a few physical demonstrations from the users, L2D2 then grounds these diverse 2D drawings in the real world. This approach reduces the time and effort required to teach the robots by enabling the users to rapidly provide a large set of diverse demonstrations. The findings from this research highlight the importance of adaptability and stability when robots and autonomous agents work around and interact with humans in diverse environments. This research contributes to the broader field of robot learning by offering scalable, adaptable, and user-friendly solutions for imitation learning and human-robot interaction, paving the way for more intuitive and robust robotic systems in human environments."],"dc:description.abstractgeneral":["As robots move beyond factory floors and start working alongside people in homes and workplaces, it becomes important for them to learn from human instructions and handle unexpected situations. One way robots can learn is by observing and copying human actions, similar to how children learn by watching adults. This approach is called imitation learning. While promising, it has challenges: robots need to handle different types of human instructions, deal with unclear or messy guidance, and keep their learning stable over time without making errors. This research focuses on creating smarter and more adaptable learning methods for robots to address these challenges. This research is organized around four main contributions. The first part of the research looks at situations where humans are available to give helpful guidance to robots. We develop a method that allows robots to learn from three types of human input: demonstrations (showing how to do a task), corrections (fixing mistakes), and preferences (choosing which robot action is better). By combining these types of input, robots can better understand what humans want and adjust their actions to complete tasks successfully. In the second part, we explore situations where human instructions are not perfect, perhaps due to the complexity of the task or physical limitations that affect their inputs. Sometimes, humans make mistakes or provide unclear feedback. To address this, we treat the robot's learning process as a system that needs to stay balanced and stable, even when the guidance is noisy. We develop a method that helps the robot understand the correct task from a wider range of human inputs without misinterpreting the inputs or becoming unstable. Next, we consider cases where a human isn't available to provide additional guidance during the robot's learning process. In such cases, robots often struggle when they face situations they haven't seen before. This can lead to the robot taking incorrect actions resulting in errors that get worse over time. To solve this problem, we create an approach called Stable Behavior Cloning, which helps robots stay consistent and perform reliably without needing additional training or complex trial-and-error learning. Finally, we think about the problem from the human user's point of view and try to make it easier and faster to teach robots. We introduce L2D2, a tool where people can teach robots by simply drawing what they want the robot to do. L2D2 uses advanced computer vision and language techniques to create different versions of the task environment automatically, so people can quickly provide many different examples. This makes teaching robots faster and less work for the user. This research shows that adaptability and stability are essential for robots to work safely and effectively alongside humans in diverse, real-world environments. The findings of this research help us move closer to building robots that are easier to use, safer, and more helpful in everyday life."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44672"],"dc:identifier.uri":["https://hdl.handle.net/10919/137800"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Imitation Learning","Human-Robot Interaction","Learning Dynamics"],"dc:title":["Robot See, Robot Do: On the Development of Robust and Adaptive Imitation Learning for Robots"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Mechanical Engineering"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:25Z"}