University of Toronto
Advancing Efficiency and Safety in Autonomous Sequential Decision Making
Abstract
dc:description.abstractThe advent of reinforcement learning (RL) has significantly transformed decision making in autonomous systems. However, its practical deployment faces substantial obstacles, chiefly in achieving sample (data) efficiency and ensuring agent safety in unpredictable, dynamic environments. Additionally, the inherent partial knowledge due to sensory and model limitations complicates agents' functionality in complex scenarios. This thesis aims to enhance sample efficiency and safety by improving adaptability in evolving environments, addressing inaccuracies under uncertainty, and implementing effective risk management strategies. It focuses on boosting RL data efficiency through simultaneous emphasis on epistemic uncertainty and knowledge transfer, and on mitigating risks associated with agent actions through aleatory uncertainty. Initially, the thesis investigates strategies to enhance RL algorithms' adaptability and knowledge transfer across varying scenarios, alongside quantifying model uncertainty—specifically, the inaccuracies in agents' understanding of their operational contexts. It underscores the importance of epistemic uncertainty within agents' approximation models, proposing a unified approach for knowledge transfer and uncertainty management to improve sample efficiency and foster the development of flexible, robust RL applications. Subsequent exploration tackles the challenges posed by partially observable Markov decision processes (POMDPs), which represent an agent's uncertainty about its environmental state. A novel framework that integrates active inference (AIF), based on the free energy principle, with RL in continuous-space POMDPs is introduced. By emphasizing the minimization of expected free energy (EFE), AIF augments RL, enhancing decision making in the face of state uncertainty. This approach notably boosts sample efficiency by systematically reducing state epistemic uncertainty, thus enabling more effective agent interaction in partially known environments. The concluding section revisits risk management, highlighting the interplay between aleatory and epistemic uncertainties in decision making. It introduces a unified uncertainty estimation algorithm that facilitates a risk-sensitive strategy, employing aleatory uncertainty while capitalizing on epistemic uncertainty to augment sample efficiency. This method presents a comprehensive strategy for navigating the complexities of risk management and enhancing sample efficiency in RL applications.
Degree
thesis:*- Department dc:contributor.department
- Electrical and Computer Engineering
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Malekzadeh, Parvin
- Advisor dc:contributor.advisor
-
- Plataniotis, Konstantinos N
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Attribution 4.0 International
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1807/140630
- OAI identifier oai:identifier
- oai:utoronto.scholaris.ca:1807/140630