{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/109429"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/109429","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Sample-efficient reinforcement learning","abstract":"Reinforcement learning has been instrumental in the recent advances made by artificial intelligence agents in various domains. Most of these advances have been abetted by the availability of huge amounts of training data. But, in several practical applications such as those arising in wireless networks, robotics, self-driving cars etc., it is expensive and sometimes completely infeasible to collect very large amounts of data. In this work, we study four different such model-free reinforcement learning problems. The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is the analysis of two-time scale reinforcement learning algorithms and the final problem we consider is the analysis of the Double Q-learning algorithm. In each of these problems, our general goal is to theoretically understand the mechanics of the different moving parts in the problem and on the basis of the insights obtained from the theory, design principled practical algorithms/heuristics that are sample-efficient.","abstract_html":"Reinforcement learning has been instrumental in the recent advances made by artificial intelligence agents in various domains. Most of these advances have been abetted by the availability of huge amounts of training data. But, in several practical applications such as those arising in wireless networks, robotics, self-driving cars etc., it is expensive and sometimes completely infeasible to collect very large amounts of data. In this work, we study four different such model-free reinforcement learning problems. The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is the analysis of two-time scale reinforcement learning algorithms and the final problem we consider is the analysis of the Double Q-learning algorithm. In each of these problems, our general goal is to theoretically understand the mechanics of the different moving parts in the problem and on the basis of the insights obtained from the theory, design principled practical algorithms/heuristics that are sample-efficient.","abstract_has_math":false,"creators":["Gupta, Harsh"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Srikant, Rayadurgam","Hajek, Bruce","Raginsky, Maxim","He, Niao"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-03-05T21:38:19Z","date_published":"2021-03-05T21:38:19Z","updated_at":"2026-07-22T22:24:50Z","subjects":["reinforcement learning","sample-efficient learning","bandits","q-learning","td-learning","stochastic approximation"],"languages":["en"],"rights":["Copyright 2020 Harsh Gupta"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/109429","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Srikant, Rayadurgam","Hajek, Bruce","Raginsky, Maxim","He, Niao"]},{"key":"dc:creator","label":"Author","values":["Gupta, Harsh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-03-05T21:38:19Z","2020-12-02","2020-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["reinforcement learning","sample-efficient learning","bandits","q-learning","td-learning","stochastic approximation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Harsh Gupta"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/109429"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Reinforcement learning has been instrumental in the recent advances made by artificial intelligence agents in various domains. Most of these advances have been abetted by the availability of huge amounts of training data. But, in several practical applications such as those arising in wireless networks, robotics, self-driving cars etc., it is expensive and sometimes completely infeasible to collect very large amounts of data. In this work, we study four different such model-free reinforcement learning problems. The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is the analysis of two-time scale reinforcement learning algorithms and the final problem we consider is the analysis of the Double Q-learning algorithm. In each of these problems, our general goal is to theoretically understand the mechanics of the different moving parts in the problem and on the basis of the insights obtained from the theory, design principled practical algorithms/heuristics that are sample-efficient.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Harsh Gupta, accepted the attached license on 2020-12-02 at 15:01.","The student, Harsh Gupta, submitted this Dissertation for approval on 2020-12-02 at 15:21.","This Dissertation was approved for publication on 2020-12-02 at 16:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16039 on 2021-03-04 at 15:36:01","Made available in DSpace on 2021-03-05T21:38:19Z (GMT). No. of bitstreams: 2 GUPTA-DISSERTATION-2020.pdf: 2634539 bytes, checksum: 1f4c4fb851cf2e0af56201f7f00c7235 (MD5) LICENSE.txt: 4208 bytes, checksum: 4c4c8c0485e1902375308c4de006a559 (MD5) Previous issue date: 2020-12-02"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Sample-efficient reinforcement learning"]}]}],"canonical_facts":{"dc:contributor":["Srikant, Rayadurgam","Hajek, Bruce","Raginsky, Maxim","He, Niao"],"dc:creator":["Gupta, Harsh"],"dc:date":["2021-03-05T21:38:19Z","2020-12-02","2020-12"],"dc:description":["Reinforcement learning has been instrumental in the recent advances made by artificial intelligence agents in various domains. Most of these advances have been abetted by the availability of huge amounts of training data. But, in several practical applications such as those arising in wireless networks, robotics, self-driving cars etc., it is expensive and sometimes completely infeasible to collect very large amounts of data. In this work, we study four different such model-free reinforcement learning problems. The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is the analysis of two-time scale reinforcement learning algorithms and the final problem we consider is the analysis of the Double Q-learning algorithm. In each of these problems, our general goal is to theoretically understand the mechanics of the different moving parts in the problem and on the basis of the insights obtained from the theory, design principled practical algorithms/heuristics that are sample-efficient.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Harsh Gupta, accepted the attached license on 2020-12-02 at 15:01.","The student, Harsh Gupta, submitted this Dissertation for approval on 2020-12-02 at 15:21.","This Dissertation was approved for publication on 2020-12-02 at 16:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16039 on 2021-03-04 at 15:36:01","Made available in DSpace on 2021-03-05T21:38:19Z (GMT). No. of bitstreams: 2 GUPTA-DISSERTATION-2020.pdf: 2634539 bytes, checksum: 1f4c4fb851cf2e0af56201f7f00c7235 (MD5) LICENSE.txt: 4208 bytes, checksum: 4c4c8c0485e1902375308c4de006a559 (MD5) Previous issue date: 2020-12-02"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/109429"],"dc:language":["en"],"dc:rights":["Copyright 2020 Harsh Gupta"],"dc:subject":["reinforcement learning","sample-efficient learning","bandits","q-learning","td-learning","stochastic approximation"],"dc:title":["Sample-efficient reinforcement learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}