University of Illinois at Urbana-Champaign
Batch value function tournament for offline policy selection in reinforcement learning
Abstract
dc:descriptionOffline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2021
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Siyuan
- Contributors dc:contributor
-
- Jiang, Nan
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2021 Siyuan Zhang
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/110602
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/110602