{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121384"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121384","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Maximum entropy on-policy reinforcement learning with monotonic policy improvement","abstract":"This Thesis was approved for publication on 2023-07-21 at 16:57.","abstract_html":"This Thesis was approved for publication on 2023-07-21 at 16:57.","abstract_has_math":false,"creators":["Kapadia, Mustafa"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Mechanical Engineering","degree_department":null,"school":null,"contributors":["Salapaka, Srinivasa M"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Entropy Maximization","Deep Reinforcement Learning","Natural Policy Gradient Methods","Combinatorial Optimization"],"languages":["en","eng"],"rights":["Copyright 2023 Mustafa Kapadia"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121384","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Salapaka, Srinivasa M"]},{"key":"dc:creator","label":"Author","values":["Kapadia, Mustafa"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-07-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechanical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Entropy Maximization","Deep Reinforcement Learning","Natural Policy Gradient Methods","Combinatorial Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Mustafa Kapadia"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121384"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This Thesis was approved for publication on 2023-07-21 at 16:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19772 on 2023-12-04 at 17:18:54","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-08-01","The student, Mustafa Kapadia, accepted the attached license on 2023-07-21 at 15:35.","The student, Mustafa Kapadia, submitted this Thesis for approval on 2023-07-21 at 15:37.","This thesis focuses on the utilization of the maximum entropy framework to train policies, which are renowned for their superior exploration and robustness, even in the presence of model and estimation errors. Our work encompasses the development of a theoretical foundation and a sample-based on-policy reinforcement learning algorithm based on the Maximum Entropy Principle (MEP). This algorithm ensures a consistent and monotonic improvement of policies across iterations, regardless of the initial policy. Furthermore, our theoretical advancements provide a framework for extending the solution of Paramterized Markov Decision Processes (ParaMDP) to address state and action spaces that were previously considered intractably large. We establish the necessary criteria for a well-posed maximum-entropy reinforcement learning problem in scenarios with an extensive number of states and actions, as well as infinite-horizon MDPs without a cost-free termination state. By incorporating the entropy over state action trajectories (or paths) into the objective function, we derive performance-estimation error bounds under MEP. This analysis involves drawing parallels and extending existing methods for on-policy reinforcement learning to cases where entropy maximization is added to the objective of the underlying optimization problem. We also introduce and analyze an ideal conservative policy iteration algorithm under MEP, and derive a practical sample-based algorithm that guarantees monotonic improvement. To evaluate the learning performance of our proposed algorithm, we conduct experiments on both continuous-control and discrete-control benchmark problems. We observe that resulting algorithms monotonic improvement with iterations and the training curve exhibits an O(1/T ) nature, where T are the number of iterations."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Maximum entropy on-policy reinforcement learning with monotonic policy improvement"]}]}],"canonical_facts":{"dc:contributor":["Salapaka, Srinivasa M"],"dc:creator":["Kapadia, Mustafa"],"dc:date":["2023-08","2023-07-21"],"dc:description":["This Thesis was approved for publication on 2023-07-21 at 16:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19772 on 2023-12-04 at 17:18:54","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-08-01","The student, Mustafa Kapadia, accepted the attached license on 2023-07-21 at 15:35.","The student, Mustafa Kapadia, submitted this Thesis for approval on 2023-07-21 at 15:37.","This thesis focuses on the utilization of the maximum entropy framework to train policies, which are renowned for their superior exploration and robustness, even in the presence of model and estimation errors. Our work encompasses the development of a theoretical foundation and a sample-based on-policy reinforcement learning algorithm based on the Maximum Entropy Principle (MEP). This algorithm ensures a consistent and monotonic improvement of policies across iterations, regardless of the initial policy. Furthermore, our theoretical advancements provide a framework for extending the solution of Paramterized Markov Decision Processes (ParaMDP) to address state and action spaces that were previously considered intractably large. We establish the necessary criteria for a well-posed maximum-entropy reinforcement learning problem in scenarios with an extensive number of states and actions, as well as infinite-horizon MDPs without a cost-free termination state. By incorporating the entropy over state action trajectories (or paths) into the objective function, we derive performance-estimation error bounds under MEP. This analysis involves drawing parallels and extending existing methods for on-policy reinforcement learning to cases where entropy maximization is added to the objective of the underlying optimization problem. We also introduce and analyze an ideal conservative policy iteration algorithm under MEP, and derive a practical sample-based algorithm that guarantees monotonic improvement. To evaluate the learning performance of our proposed algorithm, we conduct experiments on both continuous-control and discrete-control benchmark problems. We observe that resulting algorithms monotonic improvement with iterations and the training curve exhibits an O(1/T ) nature, where T are the number of iterations."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121384"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Mustafa Kapadia"],"dc:subject":["Entropy Maximization","Deep Reinforcement Learning","Natural Policy Gradient Methods","Combinatorial Optimization"],"dc:title":["Maximum entropy on-policy reinforcement learning with monotonic policy improvement"],"dc:type":["text"],"thesis:degree_discipline":["Mechanical Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}