{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108014"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108014","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Environmental curriculum learning for efficiently achieving superhuman play in games","abstract":"Reinforcement learning has made large strides in training agents to play games, including complex ones such as arcade game Pommerman and real-time strategy game StarCraft II. To allow agents to grasp the many concepts in these games, curriculum learning has been used to teach agents multiple skills over time. We present Environmental Curriculum Learning, a new technique for creating a curriculum of environment versions for an agent to learn in sequence. By adding helpful features to the state and action spaces, and then removing these helpers over the course of training, agents can focus on the fundamentals of a game one at a time. Our experiments in Pommerman illustrate the design principles of ECL, and our experiments in StarCraft II show that ECL produces agents with far better final performance than without it, when using the same training algorithm. Our StarCraft II ECL agent exceeds previous score records in a StarCraft II minigame, including human records, while taking far less training time to do so than previous approaches.","abstract_html":"Reinforcement learning has made large strides in training agents to play games, including complex ones such as arcade game Pommerman and real-time strategy game StarCraft II. To allow agents to grasp the many concepts in these games, curriculum learning has been used to teach agents multiple skills over time. We present Environmental Curriculum Learning, a new technique for creating a curriculum of environment versions for an agent to learn in sequence. By adding helpful features to the state and action spaces, and then removing these helpers over the course of training, agents can focus on the fundamentals of a game one at a time. Our experiments in Pommerman illustrate the design principles of ECL, and our experiments in StarCraft II show that ECL produces agents with far better final performance than without it, when using the same training algorithm. Our StarCraft II ECL agent exceeds previous score records in a StarCraft II minigame, including human records, while taking far less training time to do so than previous approaches.","abstract_has_math":false,"creators":["Sun, Ray"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-26T21:55:00Z","date_published":"2020-08-26T21:55:00Z","updated_at":"2026-07-22T22:24:47Z","subjects":["reinforcement learning","curriculum learning","sample efficiency","StarCraft II","Pommerman","Monte Carlo tree search"],"languages":["en"],"rights":["Copyright 2020 Ray Sun"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108014","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Sun, Ray"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-26T21:55:00Z","2020-05-11","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["reinforcement learning","curriculum learning","sample efficiency","StarCraft II","Pommerman","Monte Carlo tree search"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Ray Sun"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108014"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Reinforcement learning has made large strides in training agents to play games, including complex ones such as arcade game Pommerman and real-time strategy game StarCraft II. To allow agents to grasp the many concepts in these games, curriculum learning has been used to teach agents multiple skills over time. We present Environmental Curriculum Learning, a new technique for creating a curriculum of environment versions for an agent to learn in sequence. By adding helpful features to the state and action spaces, and then removing these helpers over the course of training, agents can focus on the fundamentals of a game one at a time. Our experiments in Pommerman illustrate the design principles of ECL, and our experiments in StarCraft II show that ECL produces agents with far better final performance than without it, when using the same training algorithm. Our StarCraft II ECL agent exceeds previous score records in a StarCraft II minigame, including human records, while taking far less training time to do so than previous approaches.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Ray Sun, accepted the attached license on 2020-05-07 at 17:27.","The student, Ray Sun, submitted this Thesis for approval on 2020-05-07 at 17:40.","This Thesis was approved for publication on 2020-05-11 at 11:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15285 on 2020-08-25 at 17:13:23","Made available in DSpace on 2020-08-26T21:55:00Z (GMT). No. of bitstreams: 32 SUN-THESIS-2020.pdf: 1897307 bytes, checksum: 4243d60bc9003da489fe439c65eb5c81 (MD5) IEEE_ECE.bst: 59476 bytes, checksum: 7668c5e97bcc2d22a9f8d4eab9b269ee (MD5) abs.tex: 1048 bytes, checksum: 13c3d6cfe85f9e12991f517a71860cee (MD5) ack.tex: 372 bytes, checksum: 6db3f73954c5a3715d5eed2c58814e74 (MD5) background.tex: 10183 bytes, checksum: ab135f4199d9d0e592c33fbdb6d3691c (MD5) buildings.png: 22573 bytes, checksum: 8f6eea5e4ff3dcb532c79ca250918d78 (MD5) csthesis.tex: 3768 bytes, checksum: b628af752d8573be3526d4f0926955a0 (MD5) discussion.tex: 3854 bytes, checksum: a411a85f284cce5de03608b449e3026f (MD5) experiments.tex: 22333 bytes, checksum: 52ca47f443aef73d4d789ff603fa9299 (MD5) intro.tex: 1784 bytes, checksum: 95a27e59a21975f2ecdadcc6e66d5037 (MD5) methods.tex: 15724 bytes, checksum: 049089992526dd29532d3fba9857f7c6 (MD5) pommerman_baseline.png: 30696 bytes, checksum: a818ca57ac7d44aea16c3a0f3707ef2e (MD5) pommerman_stage1.png: 46427 bytes, checksum: d350dbc4f10c7a68eb47ec388b03b4a1 (MD5) pommerman_stage2.png: 52041 bytes, checksum: 4a8d214dc61824d53c9ec58078d75850 (MD5) pommerman_stage3.png: 45400 bytes, checksum: b58a7d1013b5763c74b260c5e39e9024 (MD5) pommermangame.png: 57548 bytes, checksum: 568b8544e4d717e27d2d688ee8be1c0a (MD5) pommermanstart.png: 34755 bytes, checksum: 609f602ea8f272c64a306cdb06ea7bd9 (MD5) relwork.tex: 3276 bytes, checksum: 031b94c0b32194c454bc0d4c72f5d8e0 (MD5) results.tex: 17193 bytes, checksum: 0b7392d7fc941f3dc74d3153362d584f (MD5) sc2battle.png: 1214051 bytes, checksum: 4e569208adfc88140a0781b0a00d95ad (MD5) sc_baseline.png: 34407 bytes, checksum: 189d61ef81d7bf5e9457ee5fa908bd7c (MD5) sc_stage1part1.png: 36655 bytes, checksum: 018cbef725f09d51e7d38e42c8d374fd (MD5) sc_stage1part2.png: 56968 bytes, checksum: b0fe99b9a7c536947f980a86ec978c1e (MD5) sc_stage2part1.png: 38636 bytes, checksum: 8972622b32f641091ab1c0c6cdb1de59 (MD5) sc_stage2part2.png: 37357 bytes, checksum: a4432772a9b006405afcfb5203d15505 (MD5) sc_stage3.png: 36837 bytes, checksum: 7e0ccf0bdb94bd67d428693e879bcee1 (MD5) schematic.png: 18325 bytes, checksum: 91b25dec532e51e7e5ea46785eee2b19 (MD5) thesisrefs.bib: 45896 bytes, checksum: edf7f65ebdaa487ccf1e0a4503aeeece (MD5) uiuc_csthesis18.cls: 20137 bytes, checksum: e344ee8e6e216722b8f625d181a59c00 (MD5) uml1.png: 110817 bytes, checksum: 9aaea4941e0410c6a9cb1e8855e48ace (MD5) uml2.png: 36195 bytes, checksum: 2135630a4927a19ba0e9def8d2ddc926 (MD5) LICENSE.txt: 4204 bytes, checksum: 07a51fdd9f4d10540a21d09f54ad0a51 (MD5) Previous issue date: 2020-05-11"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Environmental curriculum learning for efficiently achieving superhuman play in games"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian"],"dc:creator":["Sun, Ray"],"dc:date":["2020-08-26T21:55:00Z","2020-05-11","2020-05"],"dc:description":["Reinforcement learning has made large strides in training agents to play games, including complex ones such as arcade game Pommerman and real-time strategy game StarCraft II. To allow agents to grasp the many concepts in these games, curriculum learning has been used to teach agents multiple skills over time. We present Environmental Curriculum Learning, a new technique for creating a curriculum of environment versions for an agent to learn in sequence. By adding helpful features to the state and action spaces, and then removing these helpers over the course of training, agents can focus on the fundamentals of a game one at a time. Our experiments in Pommerman illustrate the design principles of ECL, and our experiments in StarCraft II show that ECL produces agents with far better final performance than without it, when using the same training algorithm. Our StarCraft II ECL agent exceeds previous score records in a StarCraft II minigame, including human records, while taking far less training time to do so than previous approaches.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Ray Sun, accepted the attached license on 2020-05-07 at 17:27.","The student, Ray Sun, submitted this Thesis for approval on 2020-05-07 at 17:40.","This Thesis was approved for publication on 2020-05-11 at 11:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15285 on 2020-08-25 at 17:13:23","Made available in DSpace on 2020-08-26T21:55:00Z (GMT). No. of bitstreams: 32 SUN-THESIS-2020.pdf: 1897307 bytes, checksum: 4243d60bc9003da489fe439c65eb5c81 (MD5) IEEE_ECE.bst: 59476 bytes, checksum: 7668c5e97bcc2d22a9f8d4eab9b269ee (MD5) abs.tex: 1048 bytes, checksum: 13c3d6cfe85f9e12991f517a71860cee (MD5) ack.tex: 372 bytes, checksum: 6db3f73954c5a3715d5eed2c58814e74 (MD5) background.tex: 10183 bytes, checksum: ab135f4199d9d0e592c33fbdb6d3691c (MD5) buildings.png: 22573 bytes, checksum: 8f6eea5e4ff3dcb532c79ca250918d78 (MD5) csthesis.tex: 3768 bytes, checksum: b628af752d8573be3526d4f0926955a0 (MD5) discussion.tex: 3854 bytes, checksum: a411a85f284cce5de03608b449e3026f (MD5) experiments.tex: 22333 bytes, checksum: 52ca47f443aef73d4d789ff603fa9299 (MD5) intro.tex: 1784 bytes, checksum: 95a27e59a21975f2ecdadcc6e66d5037 (MD5) methods.tex: 15724 bytes, checksum: 049089992526dd29532d3fba9857f7c6 (MD5) pommerman_baseline.png: 30696 bytes, checksum: a818ca57ac7d44aea16c3a0f3707ef2e (MD5) pommerman_stage1.png: 46427 bytes, checksum: d350dbc4f10c7a68eb47ec388b03b4a1 (MD5) pommerman_stage2.png: 52041 bytes, checksum: 4a8d214dc61824d53c9ec58078d75850 (MD5) pommerman_stage3.png: 45400 bytes, checksum: b58a7d1013b5763c74b260c5e39e9024 (MD5) pommermangame.png: 57548 bytes, checksum: 568b8544e4d717e27d2d688ee8be1c0a (MD5) pommermanstart.png: 34755 bytes, checksum: 609f602ea8f272c64a306cdb06ea7bd9 (MD5) relwork.tex: 3276 bytes, checksum: 031b94c0b32194c454bc0d4c72f5d8e0 (MD5) results.tex: 17193 bytes, checksum: 0b7392d7fc941f3dc74d3153362d584f (MD5) sc2battle.png: 1214051 bytes, checksum: 4e569208adfc88140a0781b0a00d95ad (MD5) sc_baseline.png: 34407 bytes, checksum: 189d61ef81d7bf5e9457ee5fa908bd7c (MD5) sc_stage1part1.png: 36655 bytes, checksum: 018cbef725f09d51e7d38e42c8d374fd (MD5) sc_stage1part2.png: 56968 bytes, checksum: b0fe99b9a7c536947f980a86ec978c1e (MD5) sc_stage2part1.png: 38636 bytes, checksum: 8972622b32f641091ab1c0c6cdb1de59 (MD5) sc_stage2part2.png: 37357 bytes, checksum: a4432772a9b006405afcfb5203d15505 (MD5) sc_stage3.png: 36837 bytes, checksum: 7e0ccf0bdb94bd67d428693e879bcee1 (MD5) schematic.png: 18325 bytes, checksum: 91b25dec532e51e7e5ea46785eee2b19 (MD5) thesisrefs.bib: 45896 bytes, checksum: edf7f65ebdaa487ccf1e0a4503aeeece (MD5) uiuc_csthesis18.cls: 20137 bytes, checksum: e344ee8e6e216722b8f625d181a59c00 (MD5) uml1.png: 110817 bytes, checksum: 9aaea4941e0410c6a9cb1e8855e48ace (MD5) uml2.png: 36195 bytes, checksum: 2135630a4927a19ba0e9def8d2ddc926 (MD5) LICENSE.txt: 4204 bytes, checksum: 07a51fdd9f4d10540a21d09f54ad0a51 (MD5) Previous issue date: 2020-05-11"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108014"],"dc:language":["en"],"dc:rights":["Copyright 2020 Ray Sun"],"dc:subject":["reinforcement learning","curriculum learning","sample efficiency","StarCraft II","Pommerman","Monte Carlo tree search"],"dc:title":["Environmental curriculum learning for efficiently achieving superhuman play in games"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}