University of Ontario Institute of Technology
Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance
Abstract
dc:description.abstractRecent works have shown that tackling offline Reinforcement Learning (RL) with a conditional policy produces promising results. Decision Transformer (DT) have shown promising results in offline RL by leveraging sequence modeling. However, standard DTs rely on Returns-to-Go (RTG) tokens, which are heuristically defined and often suboptimal for goal-conditioned tasks. In this work, we introduce Quasimetric Decision Transformer (QuaD), a novel approach that replaces RTG with learned quasimetric distances, providing a more structured and theoretically grounded guidance signal for long-horizon decision-making. We explore two quasimetric formulations: Interval Quasimetric Embedding (IQE) and Metric Residual Network (MRN), and integrate them into DTs. Extensive evaluations on the AntMaze benchmark demonstrate that QuaD outperforms standard DTs, achieving state-of-the-art success rates and improved generalization to unseen goals. Our results suggest that quasimetric guidance is a viable alternative to RTG, opening new directions for learning structured distance representations in offline RL.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MSc)
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Ontario Institute of Technology
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Goyani, Madhav
- Advisors dc:contributor.advisor
-
- Ebrahimi, Mehran
- Davoudi, Kourosh
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10155/1960
- OAI identifier oai:identifier
- oai:ontariotechu.scholaris.ca:10155/1960