Back to results

Department of Mathematics and Applied Mathematics

Evaluating transformers as memory systems in reinforcement learning

Abstract

dc:description.abstract

Memory is an important component of effective learning systems and is crucial in non-Markovian as well as partially observable environments. In recent years, Long Short-Term Memory (LSTM) networks have been the dominant mechanism for providing memory in reinforcement learning, however, the success of transformers in natural language processing tasks has highlighted a promising and viable alternative. Memory in reinforcement learning is particularly difficult as rewards are often sparse and distributed over many time steps. Early research into transformers as memory mechanisms for reinforcement learning indicated that the canonical model is not suitable, and that additional gated recurrent units and architectural modifications are necessary to stabilize these models. Several additional improvements to the canonical model have further extended its capabilities, such as increasing the attention span, dynamically selecting the number of per-symbol processing steps and accelerating convergence. It remains unclear, however, whether combining these improvements could provide meaningful performance gains overall. This dissertation examines several extensions to the canonical Transformer as memory mechanisms in reinforcement learning and empirically studies their combination, which we term the Integrated Transformer. Our findings support prior work that suggests gating variants of the Transformer architecture may outperform LSTMs as memory networks in reinforcement learning. However, our results indicate that while gated variants of the Transformer architecture may be able to model dependencies over a longer temporal horizon, these models do not necessarily outperform LSTMs when tasked with retaining increasing quantities of information.

Degree

thesis:*
Grantor
Department of Mathematics and Applied Mathematics
Year dc:date.issued
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Makkink, Thomas
Advisors dc:contributor.advisor
  • Shock, Jonathan
  • Pretorius, Arnu

Subjects

dc:subject × 1

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/11427/35840
OAI identifier oai:identifier
oai:open.uct.ac.za:11427/35840

Chain of custody

source
Harvested from
University of Cape Town
Base URL
open.uct.ac.za/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Makkink, Thomas. Evaluating transformers as memory systems in reinforcement learning. Department of Mathematics and Applied Mathematics, 2021. http://hdl.handle.net/11427/35840