{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125792"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125792","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards simulator enabled offline validation and learning of reinforcement learning agents","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2026-08-01","abstract_has_math":false,"creators":["Katdare, Pulkit"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Driggs-Campbell, Katherine","Varshney, Lav","Schwing, Alexander","Gupta, Saurabh"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-11","date_published":"2024-07-11","updated_at":"2026-07-22T22:25:02Z","subjects":["Reinforcement Learning","Off Policy Evaluation","Offline Reinforcement Learning"],"languages":["en","eng"],"rights":["Copyright 2024 Pulkit Katdare"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125792","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Driggs-Campbell, Katherine","Varshney, Lav","Schwing, Alexander","Gupta, Saurabh"]},{"key":"dc:creator","label":"Author","values":["Katdare, Pulkit"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-11","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Off Policy Evaluation","Offline Reinforcement Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Pulkit Katdare"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125792"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Pulkit Katdare, accepted the attached license on 2024-07-09 at 13:12.","The student, Pulkit Katdare, submitted this Dissertation for approval on 2024-07-09 at 13:13.","This Dissertation was approved for publication on 2024-07-11 at 07:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20969 on 2025-02-04 at 21:25:34","Reinforcement learning pertains to a class of algorithms that attempt to learn the right decisions through rigorous self-play. Over time, reinforcement learning has achieved great success across many commercial applications like designing chat-bots, learning to play complex games and robotics. Although reinforcement learning when applied to robotics has demonstrated a lot of promise, it still has a long way to go. This is because, reinforcement learning typically requires the robot to explore diverse actions efficiently to learn the right set of actions. Although exploration is possible in domains like chess, this approach is not really viable in robotics. This is because, robots exploring their actions might be unsafe and lead to robot failures and/or accident. Commercially this can be risky, potentially involving significant sums of money. A common compromise in many reinforcement learning applications is to learn and validate robotic agents extensively in a software based simulation environment of the robot. This will not only allow robot to explore environments, but possibly validate for future safety concerns. After sufficient testing across different possible scenarios in the simulation, the idea is to deploy them to the robot. Although most common, this particular method of implementing robot learning does-not translate well in practice. This is because of the reality gap between the simulator and the real world, which can sometimes lead to drastic changes in the performance from simulation to the real world. The reality gap, which is often called as the sim2real gap, refers to the inability of the simulator to mimic each and every aspect of the robot environment like friction. Another common method to perform reinforcement learning for robots to utilize fixed batch of expert data collected on the robot. Such a method, although promising, does not work well in practice, because the robot does-not generalize well to states not seen in the offline data. In this thesis, we aim to combine different aspects of the above mentioned methods to perform offline validation and learning for robot agents. We note that simulation, although imperfect does allow for exploration and generalization for robotic agents. At the same time, offline data collected from the real world is limited but demonstrates robot's interactions in the real world which is expensive. The key idea of our thesis is to succinctly combine offline data with the simulator to both validate and learn optimal policies for the real world in simulation. Notably we look at performative correction of the simulator to ensure an accurate reflection the robot in the real world. In performative correction, we push for designing techniques which can quantitatively measure the performance of the robot in the real world. What that means is that we utilize rollouts from the simulator as before and re-weight trajectories that resemble the real world performance with a higher weight than the ones which are far off from the real world. To that end, we propose two methods to validate robot's performance in the real world. In chapter 2, we propose a technique that estimates the sim2real gap between the simulator and the real world and corrects for the gap by shaping the reward function. In chapter 3, we refine our correction by looking at the dual formulation of a reinforcement learning problem to provide a min-max optimization which is much harder but more accurate at correcting for simulator performance. In chapter 4, we finally take the first steps in building a better learning algorithm that is able to utlize these validation techniques to perform offline reinforcement learning."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards simulator enabled offline validation and learning of reinforcement learning agents"]}]}],"canonical_facts":{"dc:contributor":["Driggs-Campbell, Katherine","Varshney, Lav","Schwing, Alexander","Gupta, Saurabh"],"dc:creator":["Katdare, Pulkit"],"dc:date":["2024-07-11","2024-08"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Pulkit Katdare, accepted the attached license on 2024-07-09 at 13:12.","The student, Pulkit Katdare, submitted this Dissertation for approval on 2024-07-09 at 13:13.","This Dissertation was approved for publication on 2024-07-11 at 07:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20969 on 2025-02-04 at 21:25:34","Reinforcement learning pertains to a class of algorithms that attempt to learn the right decisions through rigorous self-play. Over time, reinforcement learning has achieved great success across many commercial applications like designing chat-bots, learning to play complex games and robotics. Although reinforcement learning when applied to robotics has demonstrated a lot of promise, it still has a long way to go. This is because, reinforcement learning typically requires the robot to explore diverse actions efficiently to learn the right set of actions. Although exploration is possible in domains like chess, this approach is not really viable in robotics. This is because, robots exploring their actions might be unsafe and lead to robot failures and/or accident. Commercially this can be risky, potentially involving significant sums of money. A common compromise in many reinforcement learning applications is to learn and validate robotic agents extensively in a software based simulation environment of the robot. This will not only allow robot to explore environments, but possibly validate for future safety concerns. After sufficient testing across different possible scenarios in the simulation, the idea is to deploy them to the robot. Although most common, this particular method of implementing robot learning does-not translate well in practice. This is because of the reality gap between the simulator and the real world, which can sometimes lead to drastic changes in the performance from simulation to the real world. The reality gap, which is often called as the sim2real gap, refers to the inability of the simulator to mimic each and every aspect of the robot environment like friction. Another common method to perform reinforcement learning for robots to utilize fixed batch of expert data collected on the robot. Such a method, although promising, does not work well in practice, because the robot does-not generalize well to states not seen in the offline data. In this thesis, we aim to combine different aspects of the above mentioned methods to perform offline validation and learning for robot agents. We note that simulation, although imperfect does allow for exploration and generalization for robotic agents. At the same time, offline data collected from the real world is limited but demonstrates robot's interactions in the real world which is expensive. The key idea of our thesis is to succinctly combine offline data with the simulator to both validate and learn optimal policies for the real world in simulation. Notably we look at performative correction of the simulator to ensure an accurate reflection the robot in the real world. In performative correction, we push for designing techniques which can quantitatively measure the performance of the robot in the real world. What that means is that we utilize rollouts from the simulator as before and re-weight trajectories that resemble the real world performance with a higher weight than the ones which are far off from the real world. To that end, we propose two methods to validate robot's performance in the real world. In chapter 2, we propose a technique that estimates the sim2real gap between the simulator and the real world and corrects for the gap by shaping the reward function. In chapter 3, we refine our correction by looking at the dual formulation of a reinforcement learning problem to provide a min-max optimization which is much harder but more accurate at correcting for simulator performance. In chapter 4, we finally take the first steps in building a better learning algorithm that is able to utlize these validation techniques to perform offline reinforcement learning."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125792"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Pulkit Katdare"],"dc:subject":["Reinforcement Learning","Off Policy Evaluation","Offline Reinforcement Learning"],"dc:title":["Towards simulator enabled offline validation and learning of reinforcement learning agents"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}