Back to results

University of Toronto

Advancing Efficiency and Safety in Autonomous Sequential Decision Making

Abstract

dc:description.abstract

The advent of reinforcement learning (RL) has significantly transformed decision making in autonomous systems. However, its practical deployment faces substantial obstacles, chiefly in achieving sample (data) efficiency and ensuring agent safety in unpredictable, dynamic environments. Additionally, the inherent partial knowledge due to sensory and model limitations complicates agents' functionality in complex scenarios. This thesis aims to enhance sample efficiency and safety by improving adaptability in evolving environments, addressing inaccuracies under uncertainty, and implementing effective risk management strategies. It focuses on boosting RL data efficiency through simultaneous emphasis on epistemic uncertainty and knowledge transfer, and on mitigating risks associated with agent actions through aleatory uncertainty. Initially, the thesis investigates strategies to enhance RL algorithms' adaptability and knowledge transfer across varying scenarios, alongside quantifying model uncertainty—specifically, the inaccuracies in agents' understanding of their operational contexts. It underscores the importance of epistemic uncertainty within agents' approximation models, proposing a unified approach for knowledge transfer and uncertainty management to improve sample efficiency and foster the development of flexible, robust RL applications. Subsequent exploration tackles the challenges posed by partially observable Markov decision processes (POMDPs), which represent an agent's uncertainty about its environmental state. A novel framework that integrates active inference (AIF), based on the free energy principle, with RL in continuous-space POMDPs is introduced. By emphasizing the minimization of expected free energy (EFE), AIF augments RL, enhancing decision making in the face of state uncertainty. This approach notably boosts sample efficiency by systematically reducing state epistemic uncertainty, thus enabling more effective agent interaction in partially known environments. The concluding section revisits risk management, highlighting the interplay between aleatory and epistemic uncertainties in decision making. It introduces a unified uncertainty estimation algorithm that facilitates a risk-sensitive strategy, employing aleatory uncertainty while capitalizing on epistemic uncertainty to augment sample efficiency. This method presents a comprehensive strategy for navigating the complexities of risk management and enhancing sample efficiency in RL applications.

Degree

thesis:*
Department dc:contributor.department
Electrical and Computer Engineering
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Malekzadeh, Parvin
Advisor dc:contributor.advisor
  • Plataniotis, Konstantinos N

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • Attribution 4.0 International

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1807/140630
OAI identifier oai:identifier
oai:utoronto.scholaris.ca:1807/140630

Chain of custody

source
Harvested from
University of Toronto
Base URL
utoronto.scholaris.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Malekzadeh, Parvin. Advancing Efficiency and Safety in Autonomous Sequential Decision Making. 2024. http://hdl.handle.net/1807/140630