University of Toronto
Who Should I Trust? Uncertainty and Risk for Knowledge Transfer from Multiple Sources in Reinforcement Learning Domains
Abstract
dc:description.abstractDespite the recent success of reinforcement learning (RL) in simulated domains and industrial applications, sample-efficiency remains a fundamental limitation of many model-free algorithms. Transfer learning mitigates this problem by using prior knowledge obtained by solving one set of tasks in order to accelerate the convergence on future tasks. However, while many frameworks have been proposed to successfully transfer different kinds of knowledge representations between tasks, existing transfer learning approaches remain largely incognizant to risk and uncertainty during transfer. In this thesis, we identify two sources of risk that must be addressed in order to make transfer learning from multiple knowledge sources more reliable and autonomous: epistemic uncertainty arises due to a lack of uncertainty about the model predictions, while aleatory uncertainty arises due to the stochastic nature of the environment. We address epistemic uncertainty by leveraging Bayesian model combination (BMC) to quantify and utilize uncertainty over the selection of knowledge sources for transfer, and we develop novel analytical techniques to efficiently tackle approximate Bayesian inference to train such models. We demonstrate the success of the proposed framework by transferring value functions, policies, and raw demonstrations between tasks. Next, we begin our treatment of aleatory uncertainty by highlighting some of the challenges in accounting for such risks during transfer, namely the lack of computationally tractable solutions that also provide theoretical assurances on the quality of transfer and control of risk. To mitigate this problem, we begin with two transfer learning approaches that have been highly successful in risk-neutral transfer -- namely potential-based reward shaping and successor features -- and extend them to the risk-sensitive setting. We empirically validate all our contributions on standard RL benchmarks, where they are shown to outperform other state-of-the-art transfer learning approaches in terms of robustness to noise and covariance shift in the training data, risk-sensitivity, and ease of interpretation.
Degree
thesis:*- Department dc:contributor.department
- Mechanical and Industrial Engineering
- Year dc:date.issued
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Gimelfarb, Michael
- Advisors dc:contributor.advisor
-
- Sanner, Scott
- Lee, Chi-Guhn
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Attribution 4.0 International
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1807/126805
- OAI identifier oai:identifier
- oai:utoronto.scholaris.ca:1807/126805