Back to results

University of Illinois Urbana-Champaign

Learning-based optimal and robust control: A policy optimization perspective

Abstract

dc:description

Reinforcement learning (RL) has achieved remarkable success in domains such as video games, the game of Go, and complex robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for tackling challenging problems involving nonlinear dynamics, high-dimensional systems, or insufficient parameterization. These qualities also make PO a promising approach for addressing control theory problems, where traditional methods often rely on well-defined models that may not be available. Despite its broad applicability, PO-based control synthesis is inherently nonconvex, posing significant challenges for achieving strong performance guarantees. This dissertation aims to address these challenges in two key areas of control theory: optimal control and robust control. These classical areas provide a solid foundation for evaluating the effectiveness of PO methods relative to traditional approaches and serve as benchmarks for studying the sample complexity of RL algorithms, which learn to control systems through repeated interactions. The first part of this dissertation focuses on two optimal control problems. The first problem addresses the linear quadratic Gaussian (LQG) control problem through a case study on a first-order single-input single-output (SISO) system using PO methods. Despite the inherent nonconvexity of the optimization landscape, we demonstrate that by appropriately parameterizing the policy class and incorporating problem-specific considerations, global convergence can be achieved in this specific scenario. This result provides a foundation for addressing the general LQG problem. Additionally, we investigate the asymptotic stability and region of attraction (ROA) of feedback control systems that utilize large-scale neural network policies and high-dimensional observations. The second part of the dissertation focuses on two cornerstone problems in robust control: structured \mathcal{H}\infty-synthesis and μ-synthesis. Both problems are characterized by nonconvex nonsmooth optimization landscapes, making their solutions particularly challenging. While advanced nonconvex nonsmooth optimization techniques have been developed for the model-based setting, much less is known about their model-free counterparts and the associated sample complexity. In this work, we extend these approaches to the model-free setting, providing theoretical guarantees on sample complexity and convergence performance. We also conduct extensive numerical studies to validate the practical utility of the proposed algorithms. Together, these contributions establish a unified framework that addresses current challenges and paves the way for a comprehensive PO theory for synthesizing dynamical systems with guaranteed stability, robustness, and optimality. Additionally, this work highlights the complexities of solving synthesis problems under partial observations and nonlinearities, offering a vision for scalable and reliable solutions to meet the demands of modern control systems.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Mechanical Engineering
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Keivan Esfahani, Darioush
Contributors dc:contributor
  • Dullerud, Geir E
  • Hu, Bin
  • Seiler, Peter J
  • West, Mathew
  • Salapaka, Srinivasa M

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Darioush Keivan Esfahani
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129357
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/129357

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Keivan Esfahani, Darioush. Learning-based optimal and robust control: A policy optimization perspective. Dissertation thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129357