Back to results

University of Illinois at Urbana-Champaign

Batch value function tournament for offline policy selection in reinforcement learning

Abstract

dc:description

Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Siyuan
Contributors dc:contributor
  • Jiang, Nan

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2021 Siyuan Zhang
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/110602
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/110602

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Zhang, Siyuan. Batch value function tournament for offline policy selection in reinforcement learning. Thesis thesis, University of Illinois at Urbana-Champaign, 2021. http://hdl.handle.net/2142/110602