{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132762"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132762","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Passivity, no-regret, and performance in online learning and games","abstract":"As autonomous AI agents become more widely deployed across dynamic, multi-agent environments, they will continuously learn and interact in real time to achieve complex goals. This thesis develops a control- and game-theoretic foundation to analyze and ultimately synthesize such systems in which adaptive agents evolve in the presence of other adaptive agents. Building on this motivation, the thesis investigates the interplay between passivity, no-regret and performance of continuous-time learning dynamics. The analysis is divided into two parts: (i) the interaction between a learning model and a dynamic, uncertain environment, and (ii) the interaction among multiple adaptive learners within a game. In the first part, the learning dynamic model is viewed as an input–output operator that maps the payoffs to strategies. Building on prior work for replicator dynamics, we show that if the learning dynamic model satisfies a passivity condition between the payoff vector and the deviation of its evolving strategy from any fixed strategy, it achieves finite regret. We then prove that this passivity condition holds for strategic higher-order variants of learning dynamics that have finite regret. We further provide numerical examples to illustrate the lack of finite regret of different evolutionary dynamic models that violate the passivity property. We also examine the fragility of the finite regret property under payoff perturbations. This raises an important question: is finite regret, by itself, a sufficient metric to assess the quality of the learning dynamic models , or should additional performance measures be considered? Motivated by this consideration, the thesis addresses the ``free-lunch'' question in no-regret learning- whether one no-regret algorithm outperform another in asymptotic average reward- so that an agent incurs regret for not having chosen a particular no-regret algorithm. We develop a control-theoretic lens in which a learning dynamic model is modeled as a cascade interconnection between a diagonal LTI map $G(s)=g(s)I_n$ and the softmax nonlinearity, linking the frequency response $g(j\\omega)$ (gain and phase) directly to asymptotic performance. We introduce payoff-based higher-order variants of replicator dynamics, anticipatory/predictive replicator dynamics, and show that the anticipatory model is dynamically equivalent to predictive replicator dynamics with a first-order low-pass predictor. An oracle (perfect-prediction) variant is proved to uniformly dominate the standard replicator dynamics, i.e., it achieves higher cumulative reward at every time horizon, across all environments. Using passivity, we cast the performance comparison as a passivity question: passivity of an associated comparison system is equivalent to uniform dominance of one learning algorithm over another. This yields several free-lunch results: predictive exponential replicator dynamics with a low-pass predictor uniformly dominates the standard exponential replicator dynamics for any payoff trajectory; moreover, any predictive replicator with a passive, asymptotically stable predictor, including anticipatory replicator dynamics, locally dominates the standard replicator. Framing the global comparison between anticipatory and standard replicator as an optimal-control problem, we show the minimal achievable performance gap is zero, implying uniform dominance of the anticipatory model across all environments. Lastly, we derive closed-form expressions for the long-run average reward and limiting strategy of replicator dynamics in arbitrary $2\\pi$-periodic environments. In the second part, the focus shifts from the interaction of a single learner with a dynamic environment to the interaction among multiple learners within a game. We establish a connection between finite regret and equilibrium-independent passivity (EI–passivity) through Best–Response Stationarity (BRS). Modeling the interaction between a learning dynamic (mapping payoffs to strategies) and a game (mapping strategies to payoffs) as a feedback interconnection, we exploit the fact that contractive games are anti–incrementally passive to show that incremental passivity is a stronger notion that implies both $\\delta$–passivity and EI–passivity. Based on this connection, we develop a passivity-based classification of learning dynamics according to the passivity notion they satisfy—namely, incremental passivity, $\\delta$–passivity, and EI–passivity—and use this classification as a framework for convergence analysis in contractive games. More generally, we develop an incremental-stability analysis for payoff-based higher-order variants of replicator dynamics in matrix contractive games. Taken together, the results of this thesis provide a unified control-theoretic framework for analyzing and comparing the performance of online learning dynamics. Beyond the theoretical significance, these results bridge control theory, online learning, and game theory, offering concepts that can guide the design of stable, efficient, and robust autonomous learning systems operating in interactive, uncertain, and multi-agent environments.","abstract_html":"As autonomous AI agents become more widely deployed across dynamic, multi-agent environments, they will continuously learn and interact in real time to achieve complex goals. This thesis develops a control- and game-theoretic foundation to analyze and ultimately synthesize such systems in which adaptive agents evolve in the presence of other adaptive agents. Building on this motivation, the thesis investigates the interplay between passivity, no-regret and performance of continuous-time learning dynamics. The analysis is divided into two parts: (i) the interaction between a learning model and a dynamic, uncertain environment, and (ii) the interaction among multiple adaptive learners within a game. In the first part, the learning dynamic model is viewed as an input–output operator that maps the payoffs to strategies. Building on prior work for replicator dynamics, we show that if the learning dynamic model satisfies a passivity condition between the payoff vector and the deviation of its evolving strategy from any fixed strategy, it achieves finite regret. We then prove that this passivity condition holds for strategic higher-order variants of learning dynamics that have finite regret. We further provide numerical examples to illustrate the lack of finite regret of different evolutionary dynamic models that violate the passivity property. We also examine the fragility of the finite regret property under payoff perturbations. This raises an important question: is finite regret, by itself, a sufficient metric to assess the quality of the learning dynamic models , or should additional performance measures be considered? Motivated by this consideration, the thesis addresses the ``free-lunch&#x27;&#x27; question in no-regret learning- whether one no-regret algorithm outperform another in asymptotic average reward- so that an agent incurs regret for not having chosen a particular no-regret algorithm. We develop a control-theoretic lens in which a learning dynamic model is modeled as a cascade interconnection between a diagonal LTI map <span class=\"etd-inline-math\">G(s)=g(s)I<sub>n</sub></span> and the softmax nonlinearity, linking the frequency response <span class=\"etd-inline-math\">g(j&omega;)</span> (gain and phase) directly to asymptotic performance. We introduce payoff-based higher-order variants of replicator dynamics, anticipatory/predictive replicator dynamics, and show that the anticipatory model is dynamically equivalent to predictive replicator dynamics with a first-order low-pass predictor. An oracle (perfect-prediction) variant is proved to uniformly dominate the standard replicator dynamics, i.e., it achieves higher cumulative reward at every time horizon, across all environments. Using passivity, we cast the performance comparison as a passivity question: passivity of an associated comparison system is equivalent to uniform dominance of one learning algorithm over another. This yields several free-lunch results: predictive exponential replicator dynamics with a low-pass predictor uniformly dominates the standard exponential replicator dynamics for any payoff trajectory; moreover, any predictive replicator with a passive, asymptotically stable predictor, including anticipatory replicator dynamics, locally dominates the standard replicator. Framing the global comparison between anticipatory and standard replicator as an optimal-control problem, we show the minimal achievable performance gap is zero, implying uniform dominance of the anticipatory model across all environments. Lastly, we derive closed-form expressions for the long-run average reward and limiting strategy of replicator dynamics in arbitrary <span class=\"etd-inline-math\">2&pi;</span>-periodic environments. In the second part, the focus shifts from the interaction of a single learner with a dynamic environment to the interaction among multiple learners within a game. We establish a connection between finite regret and equilibrium-independent passivity (EI–passivity) through Best–Response Stationarity (BRS). Modeling the interaction between a learning dynamic (mapping payoffs to strategies) and a game (mapping strategies to payoffs) as a feedback interconnection, we exploit the fact that contractive games are anti–incrementally passive to show that incremental passivity is a stronger notion that implies both <span class=\"etd-inline-math\">&delta;</span>–passivity and EI–passivity. Based on this connection, we develop a passivity-based classification of learning dynamics according to the passivity notion they satisfy—namely, incremental passivity, <span class=\"etd-inline-math\">&delta;</span>–passivity, and EI–passivity—and use this classification as a framework for convergence analysis in contractive games. More generally, we develop an incremental-stability analysis for payoff-based higher-order variants of replicator dynamics in matrix contractive games. Taken together, the results of this thesis provide a unified control-theoretic framework for analyzing and comparing the performance of online learning dynamics. Beyond the theoretical significance, these results bridge control theory, online learning, and game theory, offering concepts that can guide the design of stable, efficient, and robust autonomous learning systems operating in interactive, uncertain, and multi-agent environments.","abstract_has_math":true,"creators":["Abdelraouf, Hassan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Aerospace Engineering","degree_department":null,"school":null,"contributors":["Shamma, Jeff","Langbort, Cedric","Dullerud, Geir","Tsukamoto, Hiroyasu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Passivity No-regret Online Learning"],"languages":["en"],"rights":["© 2025 Hassan Abdelraouf All rights reserved."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132762","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shamma, Jeff","Langbort, Cedric","Dullerud, Geir","Tsukamoto, Hiroyasu"]},{"key":"dc:creator","label":"Author","values":["Abdelraouf, Hassan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-26"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Aerospace Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Passivity No-regret Online Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["© 2025 Hassan Abdelraouf All rights reserved."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132762"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As autonomous AI agents become more widely deployed across dynamic, multi-agent environments, they will continuously learn and interact in real time to achieve complex goals. This thesis develops a control- and game-theoretic foundation to analyze and ultimately synthesize such systems in which adaptive agents evolve in the presence of other adaptive agents. Building on this motivation, the thesis investigates the interplay between passivity, no-regret and performance of continuous-time learning dynamics. The analysis is divided into two parts: (i) the interaction between a learning model and a dynamic, uncertain environment, and (ii) the interaction among multiple adaptive learners within a game. In the first part, the learning dynamic model is viewed as an input–output operator that maps the payoffs to strategies. Building on prior work for replicator dynamics, we show that if the learning dynamic model satisfies a passivity condition between the payoff vector and the deviation of its evolving strategy from any fixed strategy, it achieves finite regret. We then prove that this passivity condition holds for strategic higher-order variants of learning dynamics that have finite regret. We further provide numerical examples to illustrate the lack of finite regret of different evolutionary dynamic models that violate the passivity property. We also examine the fragility of the finite regret property under payoff perturbations. This raises an important question: is finite regret, by itself, a sufficient metric to assess the quality of the learning dynamic models , or should additional performance measures be considered? Motivated by this consideration, the thesis addresses the ``free-lunch'' question in no-regret learning- whether one no-regret algorithm outperform another in asymptotic average reward- so that an agent incurs regret for not having chosen a particular no-regret algorithm. We develop a control-theoretic lens in which a learning dynamic model is modeled as a cascade interconnection between a diagonal LTI map $G(s)=g(s)I_n$ and the softmax nonlinearity, linking the frequency response $g(j\\omega)$ (gain and phase) directly to asymptotic performance. We introduce payoff-based higher-order variants of replicator dynamics, anticipatory/predictive replicator dynamics, and show that the anticipatory model is dynamically equivalent to predictive replicator dynamics with a first-order low-pass predictor. An oracle (perfect-prediction) variant is proved to uniformly dominate the standard replicator dynamics, i.e., it achieves higher cumulative reward at every time horizon, across all environments. Using passivity, we cast the performance comparison as a passivity question: passivity of an associated comparison system is equivalent to uniform dominance of one learning algorithm over another. This yields several free-lunch results: predictive exponential replicator dynamics with a low-pass predictor uniformly dominates the standard exponential replicator dynamics for any payoff trajectory; moreover, any predictive replicator with a passive, asymptotically stable predictor, including anticipatory replicator dynamics, locally dominates the standard replicator. Framing the global comparison between anticipatory and standard replicator as an optimal-control problem, we show the minimal achievable performance gap is zero, implying uniform dominance of the anticipatory model across all environments. Lastly, we derive closed-form expressions for the long-run average reward and limiting strategy of replicator dynamics in arbitrary $2\\pi$-periodic environments. In the second part, the focus shifts from the interaction of a single learner with a dynamic environment to the interaction among multiple learners within a game. We establish a connection between finite regret and equilibrium-independent passivity (EI–passivity) through Best–Response Stationarity (BRS). Modeling the interaction between a learning dynamic (mapping payoffs to strategies) and a game (mapping strategies to payoffs) as a feedback interconnection, we exploit the fact that contractive games are anti–incrementally passive to show that incremental passivity is a stronger notion that implies both $\\delta$–passivity and EI–passivity. Based on this connection, we develop a passivity-based classification of learning dynamics according to the passivity notion they satisfy—namely, incremental passivity, $\\delta$–passivity, and EI–passivity—and use this classification as a framework for convergence analysis in contractive games. More generally, we develop an incremental-stability analysis for payoff-based higher-order variants of replicator dynamics in matrix contractive games. Taken together, the results of this thesis provide a unified control-theoretic framework for analyzing and comparing the performance of online learning dynamics. Beyond the theoretical significance, these results bridge control theory, online learning, and game theory, offering concepts that can guide the design of stable, efficient, and robust autonomous learning systems operating in interactive, uncertain, and multi-agent environments.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Hassan Abdelraouf, accepted the attached license on 2025-11-24 at 11:13.","The student, Hassan Abdelraouf, submitted this Dissertation for approval on 2025-11-24 at 11:20.","This Dissertation was approved for publication on 2025-11-26 at 11:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22881 on 2026-02-19 at 20:08:47"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Passivity, no-regret, and performance in online learning and games"]}]}],"canonical_facts":{"dc:contributor":["Shamma, Jeff","Langbort, Cedric","Dullerud, Geir","Tsukamoto, Hiroyasu"],"dc:creator":["Abdelraouf, Hassan"],"dc:date":["2025-12","2025-11-26"],"dc:description":["As autonomous AI agents become more widely deployed across dynamic, multi-agent environments, they will continuously learn and interact in real time to achieve complex goals. This thesis develops a control- and game-theoretic foundation to analyze and ultimately synthesize such systems in which adaptive agents evolve in the presence of other adaptive agents. Building on this motivation, the thesis investigates the interplay between passivity, no-regret and performance of continuous-time learning dynamics. The analysis is divided into two parts: (i) the interaction between a learning model and a dynamic, uncertain environment, and (ii) the interaction among multiple adaptive learners within a game. In the first part, the learning dynamic model is viewed as an input–output operator that maps the payoffs to strategies. Building on prior work for replicator dynamics, we show that if the learning dynamic model satisfies a passivity condition between the payoff vector and the deviation of its evolving strategy from any fixed strategy, it achieves finite regret. We then prove that this passivity condition holds for strategic higher-order variants of learning dynamics that have finite regret. We further provide numerical examples to illustrate the lack of finite regret of different evolutionary dynamic models that violate the passivity property. We also examine the fragility of the finite regret property under payoff perturbations. This raises an important question: is finite regret, by itself, a sufficient metric to assess the quality of the learning dynamic models , or should additional performance measures be considered? Motivated by this consideration, the thesis addresses the ``free-lunch'' question in no-regret learning- whether one no-regret algorithm outperform another in asymptotic average reward- so that an agent incurs regret for not having chosen a particular no-regret algorithm. We develop a control-theoretic lens in which a learning dynamic model is modeled as a cascade interconnection between a diagonal LTI map $G(s)=g(s)I_n$ and the softmax nonlinearity, linking the frequency response $g(j\\omega)$ (gain and phase) directly to asymptotic performance. We introduce payoff-based higher-order variants of replicator dynamics, anticipatory/predictive replicator dynamics, and show that the anticipatory model is dynamically equivalent to predictive replicator dynamics with a first-order low-pass predictor. An oracle (perfect-prediction) variant is proved to uniformly dominate the standard replicator dynamics, i.e., it achieves higher cumulative reward at every time horizon, across all environments. Using passivity, we cast the performance comparison as a passivity question: passivity of an associated comparison system is equivalent to uniform dominance of one learning algorithm over another. This yields several free-lunch results: predictive exponential replicator dynamics with a low-pass predictor uniformly dominates the standard exponential replicator dynamics for any payoff trajectory; moreover, any predictive replicator with a passive, asymptotically stable predictor, including anticipatory replicator dynamics, locally dominates the standard replicator. Framing the global comparison between anticipatory and standard replicator as an optimal-control problem, we show the minimal achievable performance gap is zero, implying uniform dominance of the anticipatory model across all environments. Lastly, we derive closed-form expressions for the long-run average reward and limiting strategy of replicator dynamics in arbitrary $2\\pi$-periodic environments. In the second part, the focus shifts from the interaction of a single learner with a dynamic environment to the interaction among multiple learners within a game. We establish a connection between finite regret and equilibrium-independent passivity (EI–passivity) through Best–Response Stationarity (BRS). Modeling the interaction between a learning dynamic (mapping payoffs to strategies) and a game (mapping strategies to payoffs) as a feedback interconnection, we exploit the fact that contractive games are anti–incrementally passive to show that incremental passivity is a stronger notion that implies both $\\delta$–passivity and EI–passivity. Based on this connection, we develop a passivity-based classification of learning dynamics according to the passivity notion they satisfy—namely, incremental passivity, $\\delta$–passivity, and EI–passivity—and use this classification as a framework for convergence analysis in contractive games. More generally, we develop an incremental-stability analysis for payoff-based higher-order variants of replicator dynamics in matrix contractive games. Taken together, the results of this thesis provide a unified control-theoretic framework for analyzing and comparing the performance of online learning dynamics. Beyond the theoretical significance, these results bridge control theory, online learning, and game theory, offering concepts that can guide the design of stable, efficient, and robust autonomous learning systems operating in interactive, uncertain, and multi-agent environments.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Hassan Abdelraouf, accepted the attached license on 2025-11-24 at 11:13.","The student, Hassan Abdelraouf, submitted this Dissertation for approval on 2025-11-24 at 11:20.","This Dissertation was approved for publication on 2025-11-26 at 11:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22881 on 2026-02-19 at 20:08:47"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132762"],"dc:language":["en"],"dc:rights":["© 2025 Hassan Abdelraouf All rights reserved."],"dc:subject":["Passivity No-regret Online Learning"],"dc:title":["Passivity, no-regret, and performance in online learning and games"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Aerospace Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}