Abstract
dc:descriptionOptimizing sample efficiency, or the experience needed in an environment to gain satisfactory performance, is a core challenge for developing reinforcement learning agents. While imitation learning resolves this issue, it is constrained by expert performance. On the other hand, model-based strategies, which learn a world model of the environment, typically fail to approach the asymptotic performance of model-free approaches. In this thesis, we focus on combining imitation learning with model-free reinforcement learning to maximize sample efficiency and achieve higher asymptotic performance. We propose an intuitive approach to leveraging the strengths of each paradigm to produce higher rewards over a fixed number of frames when observing learned experts. We further investigate our method’s applicability to knowledge distillation for reduced-complexity agents. These studies and results lay the foundation for further study which will benefit model-free reinforcement learning as a whole.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Walia, Nikash
- Contributors dc:contributor
-
- Lazebnik, Svetlana
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Nikash Walia
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/120276