{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108132"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108132","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Robust imitation learning from observation","abstract":"Imitation learning, sometimes referred as learning from demonstrations, has been used in real world scenarios because of its sample efficiency and computational feasibility, such as autonomous driving and robotics control. However, imitation learning often suffers from compounding error and data mismatch, which leads to lack of robustness. Another drawback is that in traditional imitation learning, people usually assume that data for both states and actions is accessible. In reality, data about the action experts took may be more difficult to access than the data about state transitions. For example, a driving video clip shows each states (traffic signal, road condition, map navigation, etc.) the vehicle is in, but does not contain associated information about whether the driver steers left or right in this transition. To address these two issues, we propose an algorithm called Robust Imitation Learning from Observation (RILfO), that aims to provide robustness in an imitation learning from observation setting. First, we allow the agent to learn a policy given state-only demonstrations from experts. Second, we introduce an adversarial agent that aims to optimally destabilize the system by carefully engineering its loss function. We jointly train the agent and adversary so that the adversary is reinforced, and the agent explores more possibilities, thus becomes more robust to the various adversarial conditions. We experimentally test RILfO in multiple benchmark environments, compare RILfO with some baseline methods, demonstrate its robustness. We also discuss about its limitations and opportunities for future work.","abstract_html":"Imitation learning, sometimes referred as learning from demonstrations, has been used in real world scenarios because of its sample efficiency and computational feasibility, such as autonomous driving and robotics control. However, imitation learning often suffers from compounding error and data mismatch, which leads to lack of robustness. Another drawback is that in traditional imitation learning, people usually assume that data for both states and actions is accessible. In reality, data about the action experts took may be more difficult to access than the data about state transitions. For example, a driving video clip shows each states (traffic signal, road condition, map navigation, etc.) the vehicle is in, but does not contain associated information about whether the driver steers left or right in this transition. To address these two issues, we propose an algorithm called Robust Imitation Learning from Observation (RILfO), that aims to provide robustness in an imitation learning from observation setting. First, we allow the agent to learn a policy given state-only demonstrations from experts. Second, we introduce an adversarial agent that aims to optimally destabilize the system by carefully engineering its loss function. We jointly train the agent and adversary so that the adversary is reinforced, and the agent explores more possibilities, thus becomes more robust to the various adversarial conditions. We experimentally test RILfO in multiple benchmark environments, compare RILfO with some baseline methods, demonstrate its robustness. We also discuss about its limitations and opportunities for future work.","abstract_has_math":false,"creators":["Tang, Zhenyi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Driggs-Campbell, Katherine"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-26T23:57:19Z","date_published":"2020-08-26T23:57:19Z","updated_at":"2026-07-22T22:24:47Z","subjects":["Imitation learning, imitation learning from observation, robustness"],"languages":["en"],"rights":["Copyright 2020 Zhenyi Tang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108132","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Driggs-Campbell, Katherine"]},{"key":"dc:creator","label":"Author","values":["Tang, Zhenyi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-26T23:57:19Z","2022-08-26T23:58:55Z","2020-05-11","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Imitation learning, imitation learning from observation, robustness"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Zhenyi Tang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108132"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Imitation learning, sometimes referred as learning from demonstrations, has been used in real world scenarios because of its sample efficiency and computational feasibility, such as autonomous driving and robotics control. However, imitation learning often suffers from compounding error and data mismatch, which leads to lack of robustness. Another drawback is that in traditional imitation learning, people usually assume that data for both states and actions is accessible. In reality, data about the action experts took may be more difficult to access than the data about state transitions. For example, a driving video clip shows each states (traffic signal, road condition, map navigation, etc.) the vehicle is in, but does not contain associated information about whether the driver steers left or right in this transition. To address these two issues, we propose an algorithm called Robust Imitation Learning from Observation (RILfO), that aims to provide robustness in an imitation learning from observation setting. First, we allow the agent to learn a policy given state-only demonstrations from experts. Second, we introduce an adversarial agent that aims to optimally destabilize the system by carefully engineering its loss function. We jointly train the agent and adversary so that the adversary is reinforced, and the agent explores more possibilities, thus becomes more robust to the various adversarial conditions. We experimentally test RILfO in multiple benchmark environments, compare RILfO with some baseline methods, demonstrate its robustness. We also discuss about its limitations and opportunities for future work.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-05-01","The student, Zhenyi Tang, accepted the attached license on 2020-05-08 at 16:52.","The student, Zhenyi Tang, submitted this Thesis for approval on 2020-05-08 at 17:05.","This Thesis was approved for publication on 2020-05-11 at 11:41.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15100 on 2020-08-25 at 17:28:26","Made available in DSpace on 2020-08-26T23:57:19Z (GMT). No. of bitstreams: 2 TANG-THESIS-2020.pdf: 3361240 bytes, checksum: cb9503fba6cccd403cc5d960a1f63808 (MD5) LICENSE.txt: 4208 bytes, checksum: a1117833da9adf79708852f78e61e02c (MD5) Previous issue date: 2020-05-11","Embargo set by: Seth Robbins for item 115744 Lift date: 2022-08-26T23:57:28Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115744 Lift date: 2022-08-26T23:58:55Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Robust imitation learning from observation"]}]}],"canonical_facts":{"dc:contributor":["Driggs-Campbell, Katherine"],"dc:creator":["Tang, Zhenyi"],"dc:date":["2020-08-26T23:57:19Z","2022-08-26T23:58:55Z","2020-05-11","2020-05"],"dc:description":["Imitation learning, sometimes referred as learning from demonstrations, has been used in real world scenarios because of its sample efficiency and computational feasibility, such as autonomous driving and robotics control. However, imitation learning often suffers from compounding error and data mismatch, which leads to lack of robustness. Another drawback is that in traditional imitation learning, people usually assume that data for both states and actions is accessible. In reality, data about the action experts took may be more difficult to access than the data about state transitions. For example, a driving video clip shows each states (traffic signal, road condition, map navigation, etc.) the vehicle is in, but does not contain associated information about whether the driver steers left or right in this transition. To address these two issues, we propose an algorithm called Robust Imitation Learning from Observation (RILfO), that aims to provide robustness in an imitation learning from observation setting. First, we allow the agent to learn a policy given state-only demonstrations from experts. Second, we introduce an adversarial agent that aims to optimally destabilize the system by carefully engineering its loss function. We jointly train the agent and adversary so that the adversary is reinforced, and the agent explores more possibilities, thus becomes more robust to the various adversarial conditions. We experimentally test RILfO in multiple benchmark environments, compare RILfO with some baseline methods, demonstrate its robustness. We also discuss about its limitations and opportunities for future work.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-05-01","The student, Zhenyi Tang, accepted the attached license on 2020-05-08 at 16:52.","The student, Zhenyi Tang, submitted this Thesis for approval on 2020-05-08 at 17:05.","This Thesis was approved for publication on 2020-05-11 at 11:41.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15100 on 2020-08-25 at 17:28:26","Made available in DSpace on 2020-08-26T23:57:19Z (GMT). No. of bitstreams: 2 TANG-THESIS-2020.pdf: 3361240 bytes, checksum: cb9503fba6cccd403cc5d960a1f63808 (MD5) LICENSE.txt: 4208 bytes, checksum: a1117833da9adf79708852f78e61e02c (MD5) Previous issue date: 2020-05-11","Embargo set by: Seth Robbins for item 115744 Lift date: 2022-08-26T23:57:28Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115744 Lift date: 2022-08-26T23:58:55Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108132"],"dc:language":["en"],"dc:rights":["Copyright 2020 Zhenyi Tang"],"dc:subject":["Imitation learning, imitation learning from observation, robustness"],"dc:title":["Robust imitation learning from observation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}