{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105067"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105067","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient reinforcement learning through variance reduction and trajectory synthesis","abstract":"Reinforcement learning is a general and unified framework that has been proven promising for many important AI applications, such as robotics, self-driving vehicles. However, current reinforcement learning algorithms suffer from large variance and sampling inefficiency, which leads to slow convergent rate as well as unstable performance. In this thesis, we manage to alleviate these two relevant problems. For enormous variance, we combine variance reduced optimization with deep Q-learning. For inefficient sampling, we propose novel framework that integrates self-imitation learning and artificial synthesis procedure. Our approaches, which are flexible and could be extended to many tasks, prove their effectiveness through experiments on Atari and MuJoCo environment.","abstract_html":"Reinforcement learning is a general and unified framework that has been proven promising for many important AI applications, such as robotics, self-driving vehicles. However, current reinforcement learning algorithms suffer from large variance and sampling inefficiency, which leads to slow convergent rate as well as unstable performance. In this thesis, we manage to alleviate these two relevant problems. For enormous variance, we combine variance reduced optimization with deep Q-learning. For inefficient sampling, we propose novel framework that integrates self-imitation learning and artificial synthesis procedure. Our approaches, which are flexible and could be extended to many tasks, prove their effectiveness through experiments on Atari and MuJoCo environment.","abstract_has_math":false,"creators":["Zhao, Xiaoming"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:36:05Z","date_published":"2019-08-23T20:36:05Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Machine Learning, Reinforcement Learning, Optimization"],"languages":["en"],"rights":["Copyright 2019 Xiaoming Zhao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105067","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Zhao, Xiaoming"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:36:05Z","2021-08-24T09:15:24Z","2019-04-22","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning, Reinforcement Learning, Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Xiaoming Zhao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105067"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Reinforcement learning is a general and unified framework that has been proven promising for many important AI applications, such as robotics, self-driving vehicles. However, current reinforcement learning algorithms suffer from large variance and sampling inefficiency, which leads to slow convergent rate as well as unstable performance. In this thesis, we manage to alleviate these two relevant problems. For enormous variance, we combine variance reduced optimization with deep Q-learning. For inefficient sampling, we propose novel framework that integrates self-imitation learning and artificial synthesis procedure. Our approaches, which are flexible and could be extended to many tasks, prove their effectiveness through experiments on Atari and MuJoCo environment.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Xiaoming Zhao, accepted the attached license on 2019-04-19 at 17:57.","The student, Xiaoming Zhao, submitted this Thesis for approval on 2019-04-19 at 18:05.","This Thesis was approved for publication on 2019-04-22 at 12:05.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13800 on 2019-08-22 at 15:07:36","Made available in DSpace on 2019-08-23T20:36:05Z (GMT). No. of bitstreams: 2 ZHAO-THESIS-2019.pdf: 3129752 bytes, checksum: 893cd54892688a06910c84da70cff498 (MD5) LICENSE.txt: 4210 bytes, checksum: 0f525324228b2a83336cd4ca14d342e2 (MD5) Previous issue date: 2019-04-22","Embargo set by: Seth Robbins for item 112186 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112186 on 2021-08-24T09:15:24Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient reinforcement learning through variance reduction and trajectory synthesis"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian"],"dc:creator":["Zhao, Xiaoming"],"dc:date":["2019-08-23T20:36:05Z","2021-08-24T09:15:24Z","2019-04-22","2019-05"],"dc:description":["Reinforcement learning is a general and unified framework that has been proven promising for many important AI applications, such as robotics, self-driving vehicles. However, current reinforcement learning algorithms suffer from large variance and sampling inefficiency, which leads to slow convergent rate as well as unstable performance. In this thesis, we manage to alleviate these two relevant problems. For enormous variance, we combine variance reduced optimization with deep Q-learning. For inefficient sampling, we propose novel framework that integrates self-imitation learning and artificial synthesis procedure. Our approaches, which are flexible and could be extended to many tasks, prove their effectiveness through experiments on Atari and MuJoCo environment.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Xiaoming Zhao, accepted the attached license on 2019-04-19 at 17:57.","The student, Xiaoming Zhao, submitted this Thesis for approval on 2019-04-19 at 18:05.","This Thesis was approved for publication on 2019-04-22 at 12:05.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13800 on 2019-08-22 at 15:07:36","Made available in DSpace on 2019-08-23T20:36:05Z (GMT). No. of bitstreams: 2 ZHAO-THESIS-2019.pdf: 3129752 bytes, checksum: 893cd54892688a06910c84da70cff498 (MD5) LICENSE.txt: 4210 bytes, checksum: 0f525324228b2a83336cd4ca14d342e2 (MD5) Previous issue date: 2019-04-22","Embargo set by: Seth Robbins for item 112186 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112186 on 2021-08-24T09:15:24Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105067"],"dc:language":["en"],"dc:rights":["Copyright 2019 Xiaoming Zhao"],"dc:subject":["Machine Learning, Reinforcement Learning, Optimization"],"dc:title":["Efficient reinforcement learning through variance reduction and trajectory synthesis"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}