Back to results

ResearchSpace@Auckland

Trustworthy Reinforcement Learning under Constraints and Perturbations

Abstract

dc:description.abstract

Reinforcement Learning (RL) has demonstrated remarkable success in sequential decision making across domains such as game playing, autonomous driving, and large-scale resource allocation. However, deploying RL agents in real-world applications requires more than achieving high task performance, since it demands trustworthiness, encompassing safety, robustness, and fairness. This thesis investigates how to design RL agents that satisfy these properties in realistic, dynamic, and potentially adversarial environments. Toward this goal, we perform the following four tasks: (1) To study constrained multi-agent coordination that balances individual and collective objectives while incorporating non-reward requirements such as safety and fairness, we propose Density-Based Correlated Equilibria (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource allocation across a sequence, and a disjunctive model to represent situational constraints. (3) To enable multi-agent coordination under situational constraints, we design the Situational-Constrained DBCE (SC-DBCE) solution concept and the Situational-Constrained Correlated Policy Iteration (SC-CPI) algorithm, equipped with a violation-aware aggregation mechanism to ensure stability and convergence. (4) To enhance safety and robustness in RL under observation perturbations without relying on full system knowledge, we introduce Neural Model Predictive Shielding (NMPS), a modular shielding framework combining short-horizon trajectory prediction with real-time safety assessment. Extensive experiments across domains such as smart grids, medical and agricultural resource allocation, robotic warehouse management, and UAV navigation demonstrate that our approaches achieve both high performance and trustworthy behavior.

Degree

thesis:*
Name thesis:degree_name
PhD
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor dc:publisher
ResearchSpace@Auckland
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Libo
Advisors dc:contributor.advisor
  • Liu, Jiamou
  • Zhao, Kaiqi

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/2292/74983
OAI identifier oai:identifier
oai:researchspace.auckland.ac.nz:2292/74983

Chain of custody

source
Harvested from
University of Auckland
Base URL
researchspace.auckland.ac.nz/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Zhang, Libo. Trustworthy Reinforcement Learning under Constraints and Perturbations. Doctoral thesis, ResearchSpace@Auckland, 2025. https://hdl.handle.net/2292/74983