Back to results

University of Illinois at Urbana-Champaign

Playing Tetris with deep reinforcement learning

Abstract

dc:description

Tetris is a hard game to learn due to its random environment, large state space, and need for a long-term strategy. The offline version of the game is shown to be NP-hard. The state-of-the-art approach, CBMPI, is a hybrid algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by Stevens and Pradhan (2016) to use deep reinforcement learning to play Tetris was unsuccessful. The objective of this thesis is to explore the potential of DRL with Tetris games. We started with a baseline algorithm that uses a quadratic reward function and standard Q-learning framework. We experimented with linear-reward to discourage risky moves, harder games to reduce training time, and expected updates to reduce volatility in training. The combination of the three achieves 52-fold increase in performance over the baseline algorithm. With minimal fine-tuning and limited training time (6 hours), our final model achieves an average of 60357.7 pieces survived and an average score of 40163. Our experiments show that deep reinforcement learning has tremendous potential to create a high-performing Tetris player.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chen, Ziao
Contributors dc:contributor
  • Lu, Yi

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2021 Ziao Chen
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/110682
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/110682

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Chen, Ziao. Playing Tetris with deep reinforcement learning. Thesis thesis, University of Illinois at Urbana-Champaign, 2021. http://hdl.handle.net/2142/110682