Abstract
dc:description.abstractBreaking down large tasks into smaller sub-tasks, either to accelerate learning or to enable transfer across related environments, remains a central challenge in reinforcement learning (RL). Hierarchical Reinforcement Learning (HRL) addresses this problem by introducing temporal abstractions, often instantiated as options: temporally extended sequences of actions directed toward sub-goals. While prior work has largely focused on designing algorithms that explicitly learn such options, this work asks a different question: can options emerge naturally within standard RL frameworks? To this end, I introduce Decorrelate Cluster Temporal Activation (DCTA) Analysis, a tool for detecting option-like structures in agents that do not explicitly model them. I validate this approach on both a custom Four-Room environment and Atari benchmarks, providing evidence that naturally occurring options can be identified in conventional deep RL agents. In addition, I develop interpretation methods based on $n$-gram statistics of action sequences and mean+variance spatial mappings of agent states. These analyses show that the clusters match clear behaviours and movement patterns when the agent follows an option. Overall, the tool enables the detection of naturally emerging options in deep RL agents and, when combined with the proposed analyses, renders these options more interpretable.
Degree
thesis:*- Department dc:contributor.department
- Computing
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Khullar, Jayesh
- Advisors dc:contributor.supervisor
-
- Rivest, Francois
- Givigi, Sidney
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Attribution-NonCommercial-NoDerivatives 4.0 International
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1974/35403
- OAI identifier oai:identifier
- oai:queensu.scholaris.ca:1974/35403