{"id":{"repo_id":"usm","oai_identifier":"oai:aquila.usm.edu:masters_theses-2182"},"canonical_url":"https://search.dev.ndltd.org/etd/usm/oai:aquila.usm.edu:masters_theses-2182","repository":{"repo_id":"usm","name":"University of Southern Mississippi","base_url":"https://aquila.usm.edu/do/oai/"},"display":{"title":"Adversarial Inverse Reinforcement Learning with Noisy Observations","abstract":"<p>Inverse reinforcement learning (IRL) has emerged as a popular approach for training robots from human/expert demonstration, where a learner/robot infers the expert's hidden reward function using the demonstrations and a simulator. We argue that noise is inevitable in certain parts of the demonstration, and show that such noise does indeed deteriorate the performance of a popular and widely applied IRL method, called Adversarial IRL (AIRL). To render AIRL robust to noise, we formulate the problem of reward inference as one of log-likelihood optimization that accommodates noisy input. We adopt two techniques from the literature on learning hidden representations in sequential decision tasks and combine them with AIRL to solve this unified optimization problem. Experiments in four benchmark OpenAI Gym environments show that our proposed methods are effective in overcoming demonstration noise for the task of reward learning, but less so for the task of reproducing the expert behavior.</p>","abstract_html":"&lt;p&gt;Inverse reinforcement learning (IRL) has emerged as a popular approach for training robots from human/expert demonstration, where a learner/robot infers the expert&#x27;s hidden reward function using the demonstrations and a simulator. We argue that noise is inevitable in certain parts of the demonstration, and show that such noise does indeed deteriorate the performance of a popular and widely applied IRL method, called Adversarial IRL (AIRL). To render AIRL robust to noise, we formulate the problem of reward inference as one of log-likelihood optimization that accommodates noisy input. We adopt two techniques from the literature on learning hidden representations in sequential decision tasks and combine them with AIRL to solve this unified optimization problem. Experiments in four benchmark OpenAI Gym environments show that our proposed methods are effective in overcoming demonstration noise for the task of reward learning, but less so for the task of reproducing the expert behavior.&lt;/p&gt;","abstract_has_math":false,"creators":["Shrestha, Sagar"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Dr. Bikramjit Banerjee","Dr. Andrew Sung","Dr. Bo Li"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-01T07:00:00Z","date_published":"2025-05-01T07:00:00Z","updated_at":"2026-07-24T05:45:47Z","subjects":["Inverse Reinforcement Learning","Adversarial Inverse Reinforcement Learning","Demonstration Noise","Reward Inference","Log-Likelihood Optimization","Expert Behavior Reproduction","Artificial Intelligence and Robotics","Other Computer Engineering","Probability","Robotics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://aquila.usm.edu/masters_theses/1110","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Bikramjit Banerjee","Dr. Andrew Sung","Dr. Bo Li"]},{"key":"dc:creator","label":"Author","values":["Shrestha, Sagar"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2027-05-01T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Inverse Reinforcement Learning","Adversarial Inverse Reinforcement Learning","Demonstration Noise","Reward Inference","Log-Likelihood Optimization","Expert Behavior Reproduction","Artificial Intelligence and Robotics","Other Computer Engineering","Probability","Robotics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://aquila.usm.edu/masters_theses/1110"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Inverse reinforcement learning (IRL) has emerged as a popular approach for training robots from human/expert demonstration, where a learner/robot infers the expert's hidden reward function using the demonstrations and a simulator. We argue that noise is inevitable in certain parts of the demonstration, and show that such noise does indeed deteriorate the performance of a popular and widely applied IRL method, called Adversarial IRL (AIRL). To render AIRL robust to noise, we formulate the problem of reward inference as one of log-likelihood optimization that accommodates noisy input. We adopt two techniques from the literature on learning hidden representations in sequential decision tasks and combine them with AIRL to solve this unified optimization problem. Experiments in four benchmark OpenAI Gym environments show that our proposed methods are effective in overcoming demonstration noise for the task of reward learning, but less so for the task of reproducing the expert behavior.</p>"]},{"key":"dc:title","label":"Title","values":["Adversarial Inverse Reinforcement Learning with Noisy Observations"]}]}],"canonical_facts":{"dc:contributor":["Dr. Bikramjit Banerjee","Dr. Andrew Sung","Dr. Bo Li"],"dc:creator":["Shrestha, Sagar"],"dc:date.available":["2027-05-01T07:00:00Z"],"dc:description.abstract":["<p>Inverse reinforcement learning (IRL) has emerged as a popular approach for training robots from human/expert demonstration, where a learner/robot infers the expert's hidden reward function using the demonstrations and a simulator. We argue that noise is inevitable in certain parts of the demonstration, and show that such noise does indeed deteriorate the performance of a popular and widely applied IRL method, called Adversarial IRL (AIRL). To render AIRL robust to noise, we formulate the problem of reward inference as one of log-likelihood optimization that accommodates noisy input. We adopt two techniques from the literature on learning hidden representations in sequential decision tasks and combine them with AIRL to solve this unified optimization problem. Experiments in four benchmark OpenAI Gym environments show that our proposed methods are effective in overcoming demonstration noise for the task of reward learning, but less so for the task of reproducing the expert behavior.</p>"],"dc:identifier":["https://aquila.usm.edu/masters_theses/1110"],"dc:subject":["Inverse Reinforcement Learning","Adversarial Inverse Reinforcement Learning","Demonstration Noise","Reward Inference","Log-Likelihood Optimization","Expert Behavior Reproduction","Artificial Intelligence and Robotics","Other Computer Engineering","Probability","Robotics"],"dc:title":["Adversarial Inverse Reinforcement Learning with Noisy Observations"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:45:47Z"}