{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110602"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110602","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Batch value function tournament for offline policy selection in reinforcement learning","abstract":"Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.","abstract_html":"Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.","abstract_has_math":false,"creators":["Zhang, Siyuan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Jiang, Nan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T01:13:34Z","date_published":"2021-09-17T01:13:34Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Model Selection","Reinforcement Learning, Batch, Offline, RL"],"languages":["en"],"rights":["Copyright 2021 Siyuan Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110602","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jiang, Nan"]},{"key":"dc:creator","label":"Author","values":["Zhang, Siyuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T01:13:34Z","2021-04-29","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Model Selection","Reinforcement Learning, Batch, Offline, RL"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Siyuan Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110602"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Siyuan Zhang, accepted the attached license on 2021-04-29 at 11:35.","The student, Siyuan Zhang, submitted this Thesis for approval on 2021-04-29 at 11:53.","This Thesis was approved for publication on 2021-04-29 at 15:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16616 on 2021-09-16 at 16:49:33","Made available in DSpace on 2021-09-17T01:13:34Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2021.pdf: 5717078 bytes, checksum: ab48cdbf168751d8f34205d040d3fe22 (MD5) LICENSE.txt: 4209 bytes, checksum: e8b1baf107571a501c70618e6ecf7dd4 (MD5) Previous issue date: 2021-04-29"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Batch value function tournament for offline policy selection in reinforcement learning"]}]}],"canonical_facts":{"dc:contributor":["Jiang, Nan"],"dc:creator":["Zhang, Siyuan"],"dc:date":["2021-09-17T01:13:34Z","2021-04-29","2021-05"],"dc:description":["Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this thesis, we propose several changes to the original algorithm for adaptation to the task of offline model selection. We comprehensively experimented with the BVFT algorithm for policy selection across the various domains to evaluate and analyze the performance of BVFT. We show that BVFT achieves good performance in comparison with a number of state-of-the-art approaches. We demonstrate that BVFT is a reliable option to the problem of policy selection in offline reinforcement learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Siyuan Zhang, accepted the attached license on 2021-04-29 at 11:35.","The student, Siyuan Zhang, submitted this Thesis for approval on 2021-04-29 at 11:53.","This Thesis was approved for publication on 2021-04-29 at 15:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16616 on 2021-09-16 at 16:49:33","Made available in DSpace on 2021-09-17T01:13:34Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2021.pdf: 5717078 bytes, checksum: ab48cdbf168751d8f34205d040d3fe22 (MD5) LICENSE.txt: 4209 bytes, checksum: e8b1baf107571a501c70618e6ecf7dd4 (MD5) Previous issue date: 2021-04-29"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110602"],"dc:language":["en"],"dc:rights":["Copyright 2021 Siyuan Zhang"],"dc:subject":["Model Selection","Reinforcement Learning, Batch, Offline, RL"],"dc:title":["Batch value function tournament for offline policy selection in reinforcement learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}