{"id":{"repo_id":"auckland-ms","oai_identifier":"oai:researchspace.auckland.ac.nz:2292/74983"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-ms/oai:researchspace.auckland.ac.nz:2292/74983","repository":{"repo_id":"auckland-ms","name":"University of Auckland","base_url":"https://researchspace.auckland.ac.nz/server/oai/request"},"display":{"title":"Trustworthy Reinforcement Learning under Constraints and Perturbations","abstract":"Reinforcement Learning (RL) has demonstrated remarkable success in sequential decision making across domains such as game playing, autonomous driving, and large-scale resource allocation. However, deploying RL agents in real-world applications requires more than achieving high task performance, since it demands trustworthiness, encompassing safety, robustness, and fairness. This thesis investigates how to design RL agents that satisfy these properties in realistic, dynamic, and potentially adversarial environments. Toward this goal, we perform the following four tasks: (1) To study constrained multi-agent coordination that balances individual and collective objectives while incorporating non-reward requirements such as safety and fairness, we propose Density-Based Correlated Equilibria (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource allocation across a sequence, and a disjunctive model to represent situational constraints. (3) To enable multi-agent coordination under situational constraints, we design the Situational-Constrained DBCE (SC-DBCE) solution concept and the Situational-Constrained Correlated Policy Iteration (SC-CPI) algorithm, equipped with a violation-aware aggregation mechanism to ensure stability and convergence. (4) To enhance safety and robustness in RL under observation perturbations without relying on full system knowledge, we introduce Neural Model Predictive Shielding (NMPS), a modular shielding framework combining short-horizon trajectory prediction with real-time safety assessment. Extensive experiments across domains such as smart grids, medical and agricultural resource allocation, robotic warehouse management, and UAV navigation demonstrate that our approaches achieve both high performance and trustworthy behavior.","abstract_html":"Reinforcement Learning (RL) has demonstrated remarkable success in sequential decision making across domains such as game playing, autonomous driving, and large-scale resource allocation. However, deploying RL agents in real-world applications requires more than achieving high task performance, since it demands trustworthiness, encompassing safety, robustness, and fairness. This thesis investigates how to design RL agents that satisfy these properties in realistic, dynamic, and potentially adversarial environments. Toward this goal, we perform the following four tasks: (1) To study constrained multi-agent coordination that balances individual and collective objectives while incorporating non-reward requirements such as safety and fairness, we propose Density-Based Correlated Equilibria (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource allocation across a sequence, and a disjunctive model to represent situational constraints. (3) To enable multi-agent coordination under situational constraints, we design the Situational-Constrained DBCE (SC-DBCE) solution concept and the Situational-Constrained Correlated Policy Iteration (SC-CPI) algorithm, equipped with a violation-aware aggregation mechanism to ensure stability and convergence. (4) To enhance safety and robustness in RL under observation perturbations without relying on full system knowledge, we introduce Neural Model Predictive Shielding (NMPS), a modular shielding framework combining short-horizon trajectory prediction with real-time safety assessment. Extensive experiments across domains such as smart grids, medical and agricultural resource allocation, robotic warehouse management, and UAV navigation demonstrate that our approaches achieve both high performance and trustworthy behavior.","abstract_has_math":false,"creators":["Zhang, Libo"],"institution":"ResearchSpace@Auckland","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Liu, Jiamou","Zhao, Kaiqi"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T01:04:15Z","subjects":["Reinforcement Learning","Trustworthy AI"],"languages":[],"rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"rights_urls":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2292/74983","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Liu, Jiamou","Zhao, Kaiqi"]},{"key":"dc:creator","label":"Author","values":["Zhang, Libo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-03-09T18:41:09Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:publisher","label":"Institution","values":["ResearchSpace@Auckland"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Auckland"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Trustworthy AI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2292/74983"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Reinforcement Learning (RL) has demonstrated remarkable success in sequential decision making across domains such as game playing, autonomous driving, and large-scale resource allocation. However, deploying RL agents in real-world applications requires more than achieving high task performance, since it demands trustworthiness, encompassing safety, robustness, and fairness. This thesis investigates how to design RL agents that satisfy these properties in realistic, dynamic, and potentially adversarial environments. Toward this goal, we perform the following four tasks: (1) To study constrained multi-agent coordination that balances individual and collective objectives while incorporating non-reward requirements such as safety and fairness, we propose Density-Based Correlated Equilibria (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource allocation across a sequence, and a disjunctive model to represent situational constraints. (3) To enable multi-agent coordination under situational constraints, we design the Situational-Constrained DBCE (SC-DBCE) solution concept and the Situational-Constrained Correlated Policy Iteration (SC-CPI) algorithm, equipped with a violation-aware aggregation mechanism to ensure stability and convergence. (4) To enhance safety and robustness in RL under observation perturbations without relying on full system knowledge, we introduce Neural Model Predictive Shielding (NMPS), a modular shielding framework combining short-horizon trajectory prediction with real-time safety assessment. Extensive experiments across domains such as smart grids, medical and agricultural resource allocation, robotic warehouse management, and UAV navigation demonstrate that our approaches achieve both high performance and trustworthy behavior."]},{"key":"dc:title","label":"Title","values":["Trustworthy Reinforcement Learning under Constraints and Perturbations"]}]}],"canonical_facts":{"dc:contributor.advisor":["Liu, Jiamou","Zhao, Kaiqi"],"dc:creator":["Zhang, Libo"],"dc:date.accessioned":["2026-03-09T18:41:09Z"],"dc:date.issued":["2025"],"dc:description.abstract":["Reinforcement Learning (RL) has demonstrated remarkable success in sequential decision making across domains such as game playing, autonomous driving, and large-scale resource allocation. However, deploying RL agents in real-world applications requires more than achieving high task performance, since it demands trustworthiness, encompassing safety, robustness, and fairness. This thesis investigates how to design RL agents that satisfy these properties in realistic, dynamic, and potentially adversarial environments. Toward this goal, we perform the following four tasks: (1) To study constrained multi-agent coordination that balances individual and collective objectives while incorporating non-reward requirements such as safety and fairness, we propose Density-Based Correlated Equilibria (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource allocation across a sequence, and a disjunctive model to represent situational constraints. (3) To enable multi-agent coordination under situational constraints, we design the Situational-Constrained DBCE (SC-DBCE) solution concept and the Situational-Constrained Correlated Policy Iteration (SC-CPI) algorithm, equipped with a violation-aware aggregation mechanism to ensure stability and convergence. (4) To enhance safety and robustness in RL under observation perturbations without relying on full system knowledge, we introduce Neural Model Predictive Shielding (NMPS), a modular shielding framework combining short-horizon trajectory prediction with real-time safety assessment. Extensive experiments across domains such as smart grids, medical and agricultural resource allocation, robotic warehouse management, and UAV navigation demonstrate that our approaches achieve both high performance and trustworthy behavior."],"dc:identifier.uri":["https://hdl.handle.net/2292/74983"],"dc:publisher":["ResearchSpace@Auckland"],"dc:rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"dc:rights.uri":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"dc:subject":["Reinforcement Learning","Trustworthy AI"],"dc:title":["Trustworthy Reinforcement Learning under Constraints and Perturbations"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["The University of Auckland"]},"updated_at":"2026-07-24T01:04:15Z"}