{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129357"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129357","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning-based optimal and robust control: A policy optimization perspective","abstract":"Reinforcement learning (RL) has achieved remarkable success in domains such as video games, the game of Go, and complex robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for tackling challenging problems involving nonlinear dynamics, high-dimensional systems, or insufficient parameterization. These qualities also make PO a promising approach for addressing control theory problems, where traditional methods often rely on well-defined models that may not be available. Despite its broad applicability, PO-based control synthesis is inherently nonconvex, posing significant challenges for achieving strong performance guarantees. This dissertation aims to address these challenges in two key areas of control theory: optimal control and robust control. These classical areas provide a solid foundation for evaluating the effectiveness of PO methods relative to traditional approaches and serve as benchmarks for studying the sample complexity of RL algorithms, which learn to control systems through repeated interactions. The first part of this dissertation focuses on two optimal control problems. The first problem addresses the linear quadratic Gaussian (LQG) control problem through a case study on a first-order single-input single-output (SISO) system using PO methods. Despite the inherent nonconvexity of the optimization landscape, we demonstrate that by appropriately parameterizing the policy class and incorporating problem-specific considerations, global convergence can be achieved in this specific scenario. This result provides a foundation for addressing the general LQG problem. Additionally, we investigate the asymptotic stability and region of attraction (ROA) of feedback control systems that utilize large-scale neural network policies and high-dimensional observations. The second part of the dissertation focuses on two cornerstone problems in robust control: structured \\(\\mathcal{H}_{\\infty}\\)-synthesis and \\(\\mu\\)-synthesis. Both problems are characterized by nonconvex nonsmooth optimization landscapes, making their solutions particularly challenging. While advanced nonconvex nonsmooth optimization techniques have been developed for the model-based setting, much less is known about their model-free counterparts and the associated sample complexity. In this work, we extend these approaches to the model-free setting, providing theoretical guarantees on sample complexity and convergence performance. We also conduct extensive numerical studies to validate the practical utility of the proposed algorithms. Together, these contributions establish a unified framework that addresses current challenges and paves the way for a comprehensive PO theory for synthesizing dynamical systems with guaranteed stability, robustness, and optimality. Additionally, this work highlights the complexities of solving synthesis problems under partial observations and nonlinearities, offering a vision for scalable and reliable solutions to meet the demands of modern control systems.","abstract_html":"Reinforcement learning (RL) has achieved remarkable success in domains such as video games, the game of Go, and complex robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for tackling challenging problems involving nonlinear dynamics, high-dimensional systems, or insufficient parameterization. These qualities also make PO a promising approach for addressing control theory problems, where traditional methods often rely on well-defined models that may not be available. Despite its broad applicability, PO-based control synthesis is inherently nonconvex, posing significant challenges for achieving strong performance guarantees. This dissertation aims to address these challenges in two key areas of control theory: optimal control and robust control. These classical areas provide a solid foundation for evaluating the effectiveness of PO methods relative to traditional approaches and serve as benchmarks for studying the sample complexity of RL algorithms, which learn to control systems through repeated interactions. The first part of this dissertation focuses on two optimal control problems. The first problem addresses the linear quadratic Gaussian (LQG) control problem through a case study on a first-order single-input single-output (SISO) system using PO methods. Despite the inherent nonconvexity of the optimization landscape, we demonstrate that by appropriately parameterizing the policy class and incorporating problem-specific considerations, global convergence can be achieved in this specific scenario. This result provides a foundation for addressing the general LQG problem. Additionally, we investigate the asymptotic stability and region of attraction (ROA) of feedback control systems that utilize large-scale neural network policies and high-dimensional observations. The second part of the dissertation focuses on two cornerstone problems in robust control: structured <span class=\"etd-inline-math\">\\mathcal{H}<sub>\\infty</sub></span>-synthesis and <span class=\"etd-inline-math\">&mu;</span>-synthesis. Both problems are characterized by nonconvex nonsmooth optimization landscapes, making their solutions particularly challenging. While advanced nonconvex nonsmooth optimization techniques have been developed for the model-based setting, much less is known about their model-free counterparts and the associated sample complexity. In this work, we extend these approaches to the model-free setting, providing theoretical guarantees on sample complexity and convergence performance. We also conduct extensive numerical studies to validate the practical utility of the proposed algorithms. Together, these contributions establish a unified framework that addresses current challenges and paves the way for a comprehensive PO theory for synthesizing dynamical systems with guaranteed stability, robustness, and optimality. Additionally, this work highlights the complexities of solving synthesis problems under partial observations and nonlinearities, offering a vision for scalable and reliable solutions to meet the demands of modern control systems.","abstract_has_math":true,"creators":["Keivan Esfahani, Darioush"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Mechanical Engineering","degree_department":null,"school":null,"contributors":["Dullerud, Geir E","Hu, Bin","Seiler, Peter J","West, Mathew","Salapaka, Srinivasa M"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05","date_published":"2025-05","updated_at":"2026-07-22T22:25:05Z","subjects":["Control Theory","Policy Optimization","Reinforcement Learning"],"languages":["en"],"rights":["Copyright 2025 Darioush Keivan Esfahani"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129357","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dullerud, Geir E","Hu, Bin","Seiler, Peter J","West, Mathew","Salapaka, Srinivasa M"]},{"key":"dc:creator","label":"Author","values":["Keivan Esfahani, Darioush"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05","2024-12-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechanical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Control Theory","Policy Optimization","Reinforcement Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Darioush Keivan Esfahani"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129357"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Reinforcement learning (RL) has achieved remarkable success in domains such as video games, the game of Go, and complex robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for tackling challenging problems involving nonlinear dynamics, high-dimensional systems, or insufficient parameterization. These qualities also make PO a promising approach for addressing control theory problems, where traditional methods often rely on well-defined models that may not be available. Despite its broad applicability, PO-based control synthesis is inherently nonconvex, posing significant challenges for achieving strong performance guarantees. This dissertation aims to address these challenges in two key areas of control theory: optimal control and robust control. These classical areas provide a solid foundation for evaluating the effectiveness of PO methods relative to traditional approaches and serve as benchmarks for studying the sample complexity of RL algorithms, which learn to control systems through repeated interactions. The first part of this dissertation focuses on two optimal control problems. The first problem addresses the linear quadratic Gaussian (LQG) control problem through a case study on a first-order single-input single-output (SISO) system using PO methods. Despite the inherent nonconvexity of the optimization landscape, we demonstrate that by appropriately parameterizing the policy class and incorporating problem-specific considerations, global convergence can be achieved in this specific scenario. This result provides a foundation for addressing the general LQG problem. Additionally, we investigate the asymptotic stability and region of attraction (ROA) of feedback control systems that utilize large-scale neural network policies and high-dimensional observations. The second part of the dissertation focuses on two cornerstone problems in robust control: structured \\(\\mathcal{H}_{\\infty}\\)-synthesis and \\(\\mu\\)-synthesis. Both problems are characterized by nonconvex nonsmooth optimization landscapes, making their solutions particularly challenging. While advanced nonconvex nonsmooth optimization techniques have been developed for the model-based setting, much less is known about their model-free counterparts and the associated sample complexity. In this work, we extend these approaches to the model-free setting, providing theoretical guarantees on sample complexity and convergence performance. We also conduct extensive numerical studies to validate the practical utility of the proposed algorithms. Together, these contributions establish a unified framework that addresses current challenges and paves the way for a comprehensive PO theory for synthesizing dynamical systems with guaranteed stability, robustness, and optimality. Additionally, this work highlights the complexities of solving synthesis problems under partial observations and nonlinearities, offering a vision for scalable and reliable solutions to meet the demands of modern control systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Darioush Keivan Esfahani, accepted the attached license on 2024-12-10 at 12:45.","The student, Darioush Keivan Esfahani, submitted this Dissertation for approval on 2024-12-10 at 13:02.","This Dissertation was approved for publication on 2024-12-12 at 17:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21562 on 2025-10-19 at 18:17:04"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning-based optimal and robust control: A policy optimization perspective"]}]}],"canonical_facts":{"dc:contributor":["Dullerud, Geir E","Hu, Bin","Seiler, Peter J","West, Mathew","Salapaka, Srinivasa M"],"dc:creator":["Keivan Esfahani, Darioush"],"dc:date":["2025-05","2024-12-12"],"dc:description":["Reinforcement learning (RL) has achieved remarkable success in domains such as video games, the game of Go, and complex robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for tackling challenging problems involving nonlinear dynamics, high-dimensional systems, or insufficient parameterization. These qualities also make PO a promising approach for addressing control theory problems, where traditional methods often rely on well-defined models that may not be available. Despite its broad applicability, PO-based control synthesis is inherently nonconvex, posing significant challenges for achieving strong performance guarantees. This dissertation aims to address these challenges in two key areas of control theory: optimal control and robust control. These classical areas provide a solid foundation for evaluating the effectiveness of PO methods relative to traditional approaches and serve as benchmarks for studying the sample complexity of RL algorithms, which learn to control systems through repeated interactions. The first part of this dissertation focuses on two optimal control problems. The first problem addresses the linear quadratic Gaussian (LQG) control problem through a case study on a first-order single-input single-output (SISO) system using PO methods. Despite the inherent nonconvexity of the optimization landscape, we demonstrate that by appropriately parameterizing the policy class and incorporating problem-specific considerations, global convergence can be achieved in this specific scenario. This result provides a foundation for addressing the general LQG problem. Additionally, we investigate the asymptotic stability and region of attraction (ROA) of feedback control systems that utilize large-scale neural network policies and high-dimensional observations. The second part of the dissertation focuses on two cornerstone problems in robust control: structured \\(\\mathcal{H}_{\\infty}\\)-synthesis and \\(\\mu\\)-synthesis. Both problems are characterized by nonconvex nonsmooth optimization landscapes, making their solutions particularly challenging. While advanced nonconvex nonsmooth optimization techniques have been developed for the model-based setting, much less is known about their model-free counterparts and the associated sample complexity. In this work, we extend these approaches to the model-free setting, providing theoretical guarantees on sample complexity and convergence performance. We also conduct extensive numerical studies to validate the practical utility of the proposed algorithms. Together, these contributions establish a unified framework that addresses current challenges and paves the way for a comprehensive PO theory for synthesizing dynamical systems with guaranteed stability, robustness, and optimality. Additionally, this work highlights the complexities of solving synthesis problems under partial observations and nonlinearities, offering a vision for scalable and reliable solutions to meet the demands of modern control systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Darioush Keivan Esfahani, accepted the attached license on 2024-12-10 at 12:45.","The student, Darioush Keivan Esfahani, submitted this Dissertation for approval on 2024-12-10 at 13:02.","This Dissertation was approved for publication on 2024-12-12 at 17:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21562 on 2025-10-19 at 18:17:04"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129357"],"dc:language":["en"],"dc:rights":["Copyright 2025 Darioush Keivan Esfahani"],"dc:subject":["Control Theory","Policy Optimization","Reinforcement Learning"],"dc:title":["Learning-based optimal and robust control: A policy optimization perspective"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Mechanical Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}