{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113153"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113153","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Reinforcement learning for multi-agent and robust control systems","abstract":"Recent years have witnessed phenomenal accomplishments of reinforcement learning (RL) in many prominent sequential decision-making problems, such as playing the game of Go, playing real-time strategy games, robotic control, and autonomous driving. Motivated by these empirical successes, research toward theoretical understandings of RL algorithms has also re-gained great attention in recent years. In this dissertation, our goal is to contribute to these efforts through new approaches and tools, by developing RL algorithms for multi-agent and robust control systems, which find broad applications in the aforementioned examples and other diverse areas. In this dissertation, we consider specifically two different and fundamental settings that in general fall into the realm of RL for multi-agent and robust control systems: (i) decentralized multi-agent RL (MARL) with networked agents; (ii) H2/H∞-robust control synthesis. We develop new RL algorithms for these settings, supported by theoretical convergence guarantees. In setting i, a team of collaborative MARL agents is connected via a communication network, without the coordination of any central controller. With only neighbor-to-neighbor communications, we introduce decentralized actor-critic algorithms for each agent, and establish their convergence guarantees when linear function approximation is used. Setting ii corresponds to a classical robust control problem, with linear dynamics and robustness concerns in the H∞-norm sense. In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the controller automatically preserve a certain level of robustness stability, by following such policy search directions. This robustness-on-the-fly property is crucial for learning in safety-critical robust control systems. We then study the model-free regime, where we develop derivative-free policy gradient methods to solve the finite-horizon version of the problem, with sampled trajectories from the system and sample complexity guarantees. Interestingly, this robust control problem also unifies several other fundamental settings in control theory and game theory, including risk-sensitive linear control, i.e., linear exponential quadratic Gaussian (LEQG) control, and linear quadratic zero-sum dynamic games. The latter can be viewed as a benchmark setting for competitive multi-agent RL. Hence, our results provide policy-search methods for solving these problems unifiedly. Finally, we provide numerical results to demonstrate the computational efficiency of our policy search algorithms, compared to several existing robust control solvers.","abstract_html":"Recent years have witnessed phenomenal accomplishments of reinforcement learning (RL) in many prominent sequential decision-making problems, such as playing the game of Go, playing real-time strategy games, robotic control, and autonomous driving. Motivated by these empirical successes, research toward theoretical understandings of RL algorithms has also re-gained great attention in recent years. In this dissertation, our goal is to contribute to these efforts through new approaches and tools, by developing RL algorithms for multi-agent and robust control systems, which find broad applications in the aforementioned examples and other diverse areas. In this dissertation, we consider specifically two different and fundamental settings that in general fall into the realm of RL for multi-agent and robust control systems: (i) decentralized multi-agent RL (MARL) with networked agents; (ii) H2/H∞-robust control synthesis. We develop new RL algorithms for these settings, supported by theoretical convergence guarantees. In setting i, a team of collaborative MARL agents is connected via a communication network, without the coordination of any central controller. With only neighbor-to-neighbor communications, we introduce decentralized actor-critic algorithms for each agent, and establish their convergence guarantees when linear function approximation is used. Setting ii corresponds to a classical robust control problem, with linear dynamics and robustness concerns in the H∞-norm sense. In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the controller automatically preserve a certain level of robustness stability, by following such policy search directions. This robustness-on-the-fly property is crucial for learning in safety-critical robust control systems. We then study the model-free regime, where we develop derivative-free policy gradient methods to solve the finite-horizon version of the problem, with sampled trajectories from the system and sample complexity guarantees. Interestingly, this robust control problem also unifies several other fundamental settings in control theory and game theory, including risk-sensitive linear control, i.e., linear exponential quadratic Gaussian (LEQG) control, and linear quadratic zero-sum dynamic games. The latter can be viewed as a benchmark setting for competitive multi-agent RL. Hence, our results provide policy-search methods for solving these problems unifiedly. Finally, we provide numerical results to demonstrate the computational efficiency of our policy search algorithms, compared to several existing robust control solvers.","abstract_has_math":false,"creators":["Zhang, Kaiqing"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Basar, Tamer","Srikant, Rayadurgam","Dullerud, Geir","Raginsky, Maxim","Hu, Bin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T22:35:05Z","date_published":"2022-01-12T22:35:05Z","updated_at":"2026-07-22T22:24:53Z","subjects":["Reinforcement Learning","Multi-agent Systems","Robust Control"],"languages":["en"],"rights":["Copyright 2021 Kaiqing Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113153","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Basar, Tamer","Srikant, Rayadurgam","Dullerud, Geir","Raginsky, Maxim","Hu, Bin"]},{"key":"dc:creator","label":"Author","values":["Zhang, Kaiqing"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T22:35:05Z","2024-01-12T22:35:30Z","2021-07-08","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Multi-agent Systems","Robust Control"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Kaiqing Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113153"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Recent years have witnessed phenomenal accomplishments of reinforcement learning (RL) in many prominent sequential decision-making problems, such as playing the game of Go, playing real-time strategy games, robotic control, and autonomous driving. Motivated by these empirical successes, research toward theoretical understandings of RL algorithms has also re-gained great attention in recent years. In this dissertation, our goal is to contribute to these efforts through new approaches and tools, by developing RL algorithms for multi-agent and robust control systems, which find broad applications in the aforementioned examples and other diverse areas. In this dissertation, we consider specifically two different and fundamental settings that in general fall into the realm of RL for multi-agent and robust control systems: (i) decentralized multi-agent RL (MARL) with networked agents; (ii) H2/H∞-robust control synthesis. We develop new RL algorithms for these settings, supported by theoretical convergence guarantees. In setting i, a team of collaborative MARL agents is connected via a communication network, without the coordination of any central controller. With only neighbor-to-neighbor communications, we introduce decentralized actor-critic algorithms for each agent, and establish their convergence guarantees when linear function approximation is used. Setting ii corresponds to a classical robust control problem, with linear dynamics and robustness concerns in the H∞-norm sense. In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the controller automatically preserve a certain level of robustness stability, by following such policy search directions. This robustness-on-the-fly property is crucial for learning in safety-critical robust control systems. We then study the model-free regime, where we develop derivative-free policy gradient methods to solve the finite-horizon version of the problem, with sampled trajectories from the system and sample complexity guarantees. Interestingly, this robust control problem also unifies several other fundamental settings in control theory and game theory, including risk-sensitive linear control, i.e., linear exponential quadratic Gaussian (LEQG) control, and linear quadratic zero-sum dynamic games. The latter can be viewed as a benchmark setting for competitive multi-agent RL. Hence, our results provide policy-search methods for solving these problems unifiedly. Finally, we provide numerical results to demonstrate the computational efficiency of our policy search algorithms, compared to several existing robust control solvers.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Kaiqing Zhang, accepted the attached license on 2021-07-07 at 18:22.","The student, Kaiqing Zhang, submitted this Dissertation for approval on 2021-07-07 at 18:34.","This Dissertation was approved for publication on 2021-07-08 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16792 on 2022-01-12 at 12:53:54","Made available in DSpace on 2022-01-12T22:35:05Z (GMT). No. of bitstreams: 2 ZHANG-DISSERTATION-2021.pdf: 7919026 bytes, checksum: 558b0c38ac80f712d27931fa9027ed71 (MD5) LICENSE.txt: 4210 bytes, checksum: 1ecbf575d5bd0671dfc8e309e16e8c8a (MD5) Previous issue date: 2021-07-08","Embargo set by: Seth Robbins for item 121079 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Reinforcement learning for multi-agent and robust control systems"]}]}],"canonical_facts":{"dc:contributor":["Basar, Tamer","Srikant, Rayadurgam","Dullerud, Geir","Raginsky, Maxim","Hu, Bin"],"dc:creator":["Zhang, Kaiqing"],"dc:date":["2022-01-12T22:35:05Z","2024-01-12T22:35:30Z","2021-07-08","2021-08"],"dc:description":["Recent years have witnessed phenomenal accomplishments of reinforcement learning (RL) in many prominent sequential decision-making problems, such as playing the game of Go, playing real-time strategy games, robotic control, and autonomous driving. Motivated by these empirical successes, research toward theoretical understandings of RL algorithms has also re-gained great attention in recent years. In this dissertation, our goal is to contribute to these efforts through new approaches and tools, by developing RL algorithms for multi-agent and robust control systems, which find broad applications in the aforementioned examples and other diverse areas. In this dissertation, we consider specifically two different and fundamental settings that in general fall into the realm of RL for multi-agent and robust control systems: (i) decentralized multi-agent RL (MARL) with networked agents; (ii) H2/H∞-robust control synthesis. We develop new RL algorithms for these settings, supported by theoretical convergence guarantees. In setting i, a team of collaborative MARL agents is connected via a communication network, without the coordination of any central controller. With only neighbor-to-neighbor communications, we introduce decentralized actor-critic algorithms for each agent, and establish their convergence guarantees when linear function approximation is used. Setting ii corresponds to a classical robust control problem, with linear dynamics and robustness concerns in the H∞-norm sense. In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the controller automatically preserve a certain level of robustness stability, by following such policy search directions. This robustness-on-the-fly property is crucial for learning in safety-critical robust control systems. We then study the model-free regime, where we develop derivative-free policy gradient methods to solve the finite-horizon version of the problem, with sampled trajectories from the system and sample complexity guarantees. Interestingly, this robust control problem also unifies several other fundamental settings in control theory and game theory, including risk-sensitive linear control, i.e., linear exponential quadratic Gaussian (LEQG) control, and linear quadratic zero-sum dynamic games. The latter can be viewed as a benchmark setting for competitive multi-agent RL. Hence, our results provide policy-search methods for solving these problems unifiedly. Finally, we provide numerical results to demonstrate the computational efficiency of our policy search algorithms, compared to several existing robust control solvers.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Kaiqing Zhang, accepted the attached license on 2021-07-07 at 18:22.","The student, Kaiqing Zhang, submitted this Dissertation for approval on 2021-07-07 at 18:34.","This Dissertation was approved for publication on 2021-07-08 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16792 on 2022-01-12 at 12:53:54","Made available in DSpace on 2022-01-12T22:35:05Z (GMT). No. of bitstreams: 2 ZHANG-DISSERTATION-2021.pdf: 7919026 bytes, checksum: 558b0c38ac80f712d27931fa9027ed71 (MD5) LICENSE.txt: 4210 bytes, checksum: 1ecbf575d5bd0671dfc8e309e16e8c8a (MD5) Previous issue date: 2021-07-08","Embargo set by: Seth Robbins for item 121079 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113153"],"dc:language":["en"],"dc:rights":["Copyright 2021 Kaiqing Zhang"],"dc:subject":["Reinforcement Learning","Multi-agent Systems","Robust Control"],"dc:title":["Reinforcement learning for multi-agent and robust control systems"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}