{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113914"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113914","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Reinforcement learning with supervision beyond environmental rewards","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-04-06 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-04-06 without embargo terms","abstract_has_math":false,"creators":["Gangwani, Tanmay"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian","Forsyth, David","Zhai, ChengXiang","Gupta, Saurabh","Hofmann, Katja"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-04-29T21:34:52Z","date_published":"2022-04-29T21:34:52Z","updated_at":"2026-07-22T22:24:53Z","subjects":["Computer science"],"languages":["en","eng"],"rights":["Copyright 2021 Tanmay Gangwani"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113914","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian","Forsyth, David","Zhai, ChengXiang","Gupta, Saurabh","Hofmann, Katja"]},{"key":"dc:creator","label":"Author","values":["Gangwani, Tanmay"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-04-29T21:34:52Z","2021-12","2021-12-03"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Tanmay Gangwani"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113914"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-04-06 without embargo terms","The student, Tanmay Gangwani, accepted the attached license on 2021-12-03 at 02:45.","The student, Tanmay Gangwani, submitted this Dissertation for approval on 2021-12-03 at 02:56.","This Dissertation was approved for publication on 2021-12-03 at 08:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17380 on 2022-04-06 at 17:11:00","Made available in DSpace on 2022-04-29T21:34:52Z (GMT). No. of bitstreams: 3 GANGWANI-DISSERTATION-2021.pdf: 9177015 bytes, checksum: a7df3c233f14719e2fa6b7fa7cd24437 (MD5) LICENSE.txt: 4212 bytes, checksum: a623bd97d61d715ad219652cce50dddc (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: ea057c5b4deb29cb1f27e133a3a25bed (MD5) Previous issue date: 2021-12-03","Reinforcement Learning (RL) is an elegant approach to tackle sequential decision-making problems. In the standard setting, the task designer curates a reward function and the RL agent's objective is to take actions in the environment such that the long-term cumulative reward is maximized. Deep RL algorithms---that combine RL principles with deep neural networks---have been successfully used to learn behaviors in complex environments but are generally quite sensitive to the nature of the reward function. For a given RL problem, the environmental rewards could be sparse, delayed, misspecified, or unavailable (i.e., impossible to define mathematically for the required behavior). These scenarios exacerbate the challenge of training a stable deep-RL agent in a sample-efficient manner. In this thesis, we study methods that go beyond a direct reliance on the environmental rewards by generating additional information signals that the RL agent could incorporate for learning the desired skills. We start by investigating the performance bottlenecks in delayed reward environments and propose to address these by learning surrogate rewards. We include two methods to compute the surrogate rewards using the agent-environment interaction data. Then, we consider the imitation-learning (IL) setting where we don't have access to any rewards, but instead, are provided with a dataset of expert demonstrations that the RL agent must learn to reliably reproduce. We propose IL algorithms for partially observable environments and situations with discrepancies between the transition dynamics of the expert and the imitator. Next, we consider the benefits of learning an ensemble of RL agents with explicit diversity pressure. We show that diversity encourages exploration and facilitates the discovery of sparse environmental rewards. Finally, we analyze the concept of sharing knowledge between RL agents operating in different but related environments and show that the information transfer can accelerate learning."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Reinforcement learning with supervision beyond environmental rewards"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian","Forsyth, David","Zhai, ChengXiang","Gupta, Saurabh","Hofmann, Katja"],"dc:creator":["Gangwani, Tanmay"],"dc:date":["2022-04-29T21:34:52Z","2021-12","2021-12-03"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-04-06 without embargo terms","The student, Tanmay Gangwani, accepted the attached license on 2021-12-03 at 02:45.","The student, Tanmay Gangwani, submitted this Dissertation for approval on 2021-12-03 at 02:56.","This Dissertation was approved for publication on 2021-12-03 at 08:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17380 on 2022-04-06 at 17:11:00","Made available in DSpace on 2022-04-29T21:34:52Z (GMT). No. of bitstreams: 3 GANGWANI-DISSERTATION-2021.pdf: 9177015 bytes, checksum: a7df3c233f14719e2fa6b7fa7cd24437 (MD5) LICENSE.txt: 4212 bytes, checksum: a623bd97d61d715ad219652cce50dddc (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: ea057c5b4deb29cb1f27e133a3a25bed (MD5) Previous issue date: 2021-12-03","Reinforcement Learning (RL) is an elegant approach to tackle sequential decision-making problems. In the standard setting, the task designer curates a reward function and the RL agent's objective is to take actions in the environment such that the long-term cumulative reward is maximized. Deep RL algorithms---that combine RL principles with deep neural networks---have been successfully used to learn behaviors in complex environments but are generally quite sensitive to the nature of the reward function. For a given RL problem, the environmental rewards could be sparse, delayed, misspecified, or unavailable (i.e., impossible to define mathematically for the required behavior). These scenarios exacerbate the challenge of training a stable deep-RL agent in a sample-efficient manner. In this thesis, we study methods that go beyond a direct reliance on the environmental rewards by generating additional information signals that the RL agent could incorporate for learning the desired skills. We start by investigating the performance bottlenecks in delayed reward environments and propose to address these by learning surrogate rewards. We include two methods to compute the surrogate rewards using the agent-environment interaction data. Then, we consider the imitation-learning (IL) setting where we don't have access to any rewards, but instead, are provided with a dataset of expert demonstrations that the RL agent must learn to reliably reproduce. We propose IL algorithms for partially observable environments and situations with discrepancies between the transition dynamics of the expert and the imitator. Next, we consider the benefits of learning an ensemble of RL agents with explicit diversity pressure. We show that diversity encourages exploration and facilitates the discovery of sparse environmental rewards. Finally, we analyze the concept of sharing knowledge between RL agents operating in different but related environments and show that the information transfer can accelerate learning."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113914"],"dc:language":["en","eng"],"dc:rights":["Copyright 2021 Tanmay Gangwani"],"dc:subject":["Computer science"],"dc:title":["Reinforcement learning with supervision beyond environmental rewards"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}