{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110682"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110682","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Playing Tetris with deep reinforcement learning","abstract":"Tetris is a hard game to learn due to its random environment, large state space, and need for a long-term strategy. The offline version of the game is shown to be NP-hard. The state-of-the-art approach, CBMPI, is a hybrid algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by Stevens and Pradhan (2016) to use deep reinforcement learning to play Tetris was unsuccessful. The objective of this thesis is to explore the potential of DRL with Tetris games. We started with a baseline algorithm that uses a quadratic reward function and standard Q-learning framework. We experimented with linear-reward to discourage risky moves, harder games to reduce training time, and expected updates to reduce volatility in training. The combination of the three achieves 52-fold increase in performance over the baseline algorithm. With minimal fine-tuning and limited training time (6 hours), our final model achieves an average of 60357.7 pieces survived and an average score of 40163. Our experiments show that deep reinforcement learning has tremendous potential to create a high-performing Tetris player.","abstract_html":"Tetris is a hard game to learn due to its random environment, large state space, and need for a long-term strategy. The offline version of the game is shown to be NP-hard. The state-of-the-art approach, CBMPI, is a hybrid algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by Stevens and Pradhan (2016) to use deep reinforcement learning to play Tetris was unsuccessful. The objective of this thesis is to explore the potential of DRL with Tetris games. We started with a baseline algorithm that uses a quadratic reward function and standard Q-learning framework. We experimented with linear-reward to discourage risky moves, harder games to reduce training time, and expected updates to reduce volatility in training. The combination of the three achieves 52-fold increase in performance over the baseline algorithm. With minimal fine-tuning and limited training time (6 hours), our final model achieves an average of 60357.7 pieces survived and an average score of 40163. Our experiments show that deep reinforcement learning has tremendous potential to create a high-performing Tetris player.","abstract_has_math":false,"creators":["Chen, Ziao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Lu, Yi"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T02:34:33Z","date_published":"2021-09-17T02:34:33Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Reinforcement learning","Tetris","Q-learning"],"languages":["en"],"rights":["Copyright 2021 Ziao Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110682","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lu, Yi"]},{"key":"dc:creator","label":"Author","values":["Chen, Ziao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T02:34:33Z","2023-09-17T02:34:57Z","2021-04-16","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement learning","Tetris","Q-learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Ziao Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110682"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Tetris is a hard game to learn due to its random environment, large state space, and need for a long-term strategy. The offline version of the game is shown to be NP-hard. The state-of-the-art approach, CBMPI, is a hybrid algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by Stevens and Pradhan (2016) to use deep reinforcement learning to play Tetris was unsuccessful. The objective of this thesis is to explore the potential of DRL with Tetris games. We started with a baseline algorithm that uses a quadratic reward function and standard Q-learning framework. We experimented with linear-reward to discourage risky moves, harder games to reduce training time, and expected updates to reduce volatility in training. The combination of the three achieves 52-fold increase in performance over the baseline algorithm. With minimal fine-tuning and limited training time (6 hours), our final model achieves an average of 60357.7 pieces survived and an average score of 40163. Our experiments show that deep reinforcement learning has tremendous potential to create a high-performing Tetris player.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-05-01","The student, Ziao Chen, accepted the attached license on 2021-04-16 at 09:20.","The student, Ziao Chen, submitted this Thesis for approval on 2021-04-16 at 09:40.","This Thesis was approved for publication on 2021-04-16 at 13:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16364 on 2021-09-16 at 17:03:30","Made available in DSpace on 2021-09-17T02:34:33Z (GMT). No. of bitstreams: 2 CHEN-THESIS-2021.pdf: 371389 bytes, checksum: 3dfe3d92c462ddd774cb0ff52a026ec9 (MD5) LICENSE.txt: 4206 bytes, checksum: a92f53d3f04923d70b90ac96954ba69d (MD5) Previous issue date: 2021-04-16","Embargo set by: Seth Robbins for item 118525 Lift date: 2023-09-17T02:34:57Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Playing Tetris with deep reinforcement learning"]}]}],"canonical_facts":{"dc:contributor":["Lu, Yi"],"dc:creator":["Chen, Ziao"],"dc:date":["2021-09-17T02:34:33Z","2023-09-17T02:34:57Z","2021-04-16","2021-05"],"dc:description":["Tetris is a hard game to learn due to its random environment, large state space, and need for a long-term strategy. The offline version of the game is shown to be NP-hard. The state-of-the-art approach, CBMPI, is a hybrid algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by Stevens and Pradhan (2016) to use deep reinforcement learning to play Tetris was unsuccessful. The objective of this thesis is to explore the potential of DRL with Tetris games. We started with a baseline algorithm that uses a quadratic reward function and standard Q-learning framework. We experimented with linear-reward to discourage risky moves, harder games to reduce training time, and expected updates to reduce volatility in training. The combination of the three achieves 52-fold increase in performance over the baseline algorithm. With minimal fine-tuning and limited training time (6 hours), our final model achieves an average of 60357.7 pieces survived and an average score of 40163. Our experiments show that deep reinforcement learning has tremendous potential to create a high-performing Tetris player.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-05-01","The student, Ziao Chen, accepted the attached license on 2021-04-16 at 09:20.","The student, Ziao Chen, submitted this Thesis for approval on 2021-04-16 at 09:40.","This Thesis was approved for publication on 2021-04-16 at 13:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16364 on 2021-09-16 at 17:03:30","Made available in DSpace on 2021-09-17T02:34:33Z (GMT). No. of bitstreams: 2 CHEN-THESIS-2021.pdf: 371389 bytes, checksum: 3dfe3d92c462ddd774cb0ff52a026ec9 (MD5) LICENSE.txt: 4206 bytes, checksum: a92f53d3f04923d70b90ac96954ba69d (MD5) Previous issue date: 2021-04-16","Embargo set by: Seth Robbins for item 118525 Lift date: 2023-09-17T02:34:57Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110682"],"dc:language":["en"],"dc:rights":["Copyright 2021 Ziao Chen"],"dc:subject":["Reinforcement learning","Tetris","Q-learning"],"dc:title":["Playing Tetris with deep reinforcement learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}