{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/99228"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/99228","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Faster apprenticeship learning through inverse optimal control","abstract":"One of the fundamental problems of artificial intelligence is learning how to behave optimally. With applications ranging from self-driving cars to medical devices, this task is vital to modern society. There are two complementary problems in this area – reinforcement learning and inverse reinforcement learning. While reinforcement learning tries to find an optimal strategy in a given environment with known rewards for each action, inverse reinforcement learning or inverse optimal control seeks to recover rewards associated with actions given the environment and an optimal policy. Typically, apprenticeship learning is approached as a combination of these two techniques. This is an iterative process – at each step inverse reinforcement learning is applied first to get the rewards, followed by reinforcement learning to produce a guess for an optimal policy. Each guess is used in the further iterations to come up with a more accurate estimate of the reward function. While this works for problems with a small number of discreet states, the approach scales poorly. In order to mitigate those limitations, this research proposes a robust approach based on recent advances in the field of deep learning. Using the matrix formulation of inverse reinforcement learning, a reward function and an optimal policy can be recovered without having to iteratively optimize both. The approach scales well for problems with very large and continuous state spaces such as autonomous vehicle navigation. An evaluation performed using OpenAI RLLab suggests that this method is robust and ready to be adopted for solving problems both in research and various industries.","abstract_html":"One of the fundamental problems of artificial intelligence is learning how to behave optimally. With applications ranging from self-driving cars to medical devices, this task is vital to modern society. There are two complementary problems in this area – reinforcement learning and inverse reinforcement learning. While reinforcement learning tries to find an optimal strategy in a given environment with known rewards for each action, inverse reinforcement learning or inverse optimal control seeks to recover rewards associated with actions given the environment and an optimal policy. Typically, apprenticeship learning is approached as a combination of these two techniques. This is an iterative process – at each step inverse reinforcement learning is applied first to get the rewards, followed by reinforcement learning to produce a guess for an optimal policy. Each guess is used in the further iterations to come up with a more accurate estimate of the reward function. While this works for problems with a small number of discreet states, the approach scales poorly. In order to mitigate those limitations, this research proposes a robust approach based on recent advances in the field of deep learning. Using the matrix formulation of inverse reinforcement learning, a reward function and an optimal policy can be recovered without having to iteratively optimize both. The approach scales well for problems with very large and continuous state spaces such as autonomous vehicle navigation. An evaluation performed using OpenAI RLLab suggests that this method is robust and ready to be adopted for solving problems both in research and various industries.","abstract_has_math":false,"creators":["Zaytsev, Andrey"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-03-13T15:25:25Z","date_published":"2018-03-13T15:25:25Z","updated_at":"2026-07-22T22:24:37Z","subjects":["Apprenticeship learning","Inverse reinforcement learning","Inverse optimal control","Deep learning","Reinforcement learning","Machine learning"],"languages":["en"],"rights":["Copyright 2017 Andrey Zaytsev"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/99228","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Zaytsev, Andrey"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-03-13T15:25:25Z","2020-03-14T09:15:25Z","2017-12-05","2017-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Apprenticeship learning","Inverse reinforcement learning","Inverse optimal control","Deep learning","Reinforcement learning","Machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Andrey Zaytsev"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/99228"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["One of the fundamental problems of artificial intelligence is learning how to behave optimally. With applications ranging from self-driving cars to medical devices, this task is vital to modern society. There are two complementary problems in this area – reinforcement learning and inverse reinforcement learning. While reinforcement learning tries to find an optimal strategy in a given environment with known rewards for each action, inverse reinforcement learning or inverse optimal control seeks to recover rewards associated with actions given the environment and an optimal policy. Typically, apprenticeship learning is approached as a combination of these two techniques. This is an iterative process – at each step inverse reinforcement learning is applied first to get the rewards, followed by reinforcement learning to produce a guess for an optimal policy. Each guess is used in the further iterations to come up with a more accurate estimate of the reward function. While this works for problems with a small number of discreet states, the approach scales poorly. In order to mitigate those limitations, this research proposes a robust approach based on recent advances in the field of deep learning. Using the matrix formulation of inverse reinforcement learning, a reward function and an optimal policy can be recovered without having to iteratively optimize both. The approach scales well for problems with very large and continuous state spaces such as autonomous vehicle navigation. An evaluation performed using OpenAI RLLab suggests that this method is robust and ready to be adopted for solving problems both in research and various industries.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-12-01","The student, Andrey Zaytsev, accepted the attached license on 2017-12-04 at 14:08.","The student, Andrey Zaytsev, submitted this Thesis for approval on 2017-12-04 at 14:16.","This Thesis was approved for publication on 2017-12-05 at 11:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11831 on 2018-03-13 at 09:56:47","Made available in DSpace on 2018-03-13T15:25:25Z (GMT). No. of bitstreams: 2 ZAYTSEV-THESIS-2017.pdf: 991863 bytes, checksum: f8a9510579b2d52ec4d664069f440794 (MD5) LICENSE.txt: 4211 bytes, checksum: eafdd41d1297a2f23f499c286b573eec (MD5) Previous issue date: 2017-12-05","Embargo set by: Seth Robbins for item 105191 Lift date: 2020-03-13T15:25:40Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105191 Lift date: 2020-03-13T15:28:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105191 on 2020-03-14T09:15:25Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Faster apprenticeship learning through inverse optimal control"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian"],"dc:creator":["Zaytsev, Andrey"],"dc:date":["2018-03-13T15:25:25Z","2020-03-14T09:15:25Z","2017-12-05","2017-12"],"dc:description":["One of the fundamental problems of artificial intelligence is learning how to behave optimally. With applications ranging from self-driving cars to medical devices, this task is vital to modern society. There are two complementary problems in this area – reinforcement learning and inverse reinforcement learning. While reinforcement learning tries to find an optimal strategy in a given environment with known rewards for each action, inverse reinforcement learning or inverse optimal control seeks to recover rewards associated with actions given the environment and an optimal policy. Typically, apprenticeship learning is approached as a combination of these two techniques. This is an iterative process – at each step inverse reinforcement learning is applied first to get the rewards, followed by reinforcement learning to produce a guess for an optimal policy. Each guess is used in the further iterations to come up with a more accurate estimate of the reward function. While this works for problems with a small number of discreet states, the approach scales poorly. In order to mitigate those limitations, this research proposes a robust approach based on recent advances in the field of deep learning. Using the matrix formulation of inverse reinforcement learning, a reward function and an optimal policy can be recovered without having to iteratively optimize both. The approach scales well for problems with very large and continuous state spaces such as autonomous vehicle navigation. An evaluation performed using OpenAI RLLab suggests that this method is robust and ready to be adopted for solving problems both in research and various industries.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-12-01","The student, Andrey Zaytsev, accepted the attached license on 2017-12-04 at 14:08.","The student, Andrey Zaytsev, submitted this Thesis for approval on 2017-12-04 at 14:16.","This Thesis was approved for publication on 2017-12-05 at 11:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11831 on 2018-03-13 at 09:56:47","Made available in DSpace on 2018-03-13T15:25:25Z (GMT). No. of bitstreams: 2 ZAYTSEV-THESIS-2017.pdf: 991863 bytes, checksum: f8a9510579b2d52ec4d664069f440794 (MD5) LICENSE.txt: 4211 bytes, checksum: eafdd41d1297a2f23f499c286b573eec (MD5) Previous issue date: 2017-12-05","Embargo set by: Seth Robbins for item 105191 Lift date: 2020-03-13T15:25:40Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105191 Lift date: 2020-03-13T15:28:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105191 on 2020-03-14T09:15:25Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/99228"],"dc:language":["en"],"dc:rights":["Copyright 2017 Andrey Zaytsev"],"dc:subject":["Apprenticeship learning","Inverse reinforcement learning","Inverse optimal control","Deep learning","Reinforcement learning","Machine learning"],"dc:title":["Faster apprenticeship learning through inverse optimal control"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:37Z"}