Back to results

University of Ontario Institute of Technology

Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance

Abstract

dc:description.abstract

Recent works have shown that tackling offline Reinforcement Learning (RL) with a conditional policy produces promising results. Decision Transformer (DT) have shown promising results in offline RL by leveraging sequence modeling. However, standard DTs rely on Returns-to-Go (RTG) tokens, which are heuristically defined and often suboptimal for goal-conditioned tasks. In this work, we introduce Quasimetric Decision Transformer (QuaD), a novel approach that replaces RTG with learned quasimetric distances, providing a more structured and theoretically grounded guidance signal for long-horizon decision-making. We explore two quasimetric formulations: Interval Quasimetric Embedding (IQE) and Metric Residual Network (MRN), and integrate them into DTs. Extensive evaluations on the AntMaze benchmark demonstrate that QuaD outperforms standard DTs, achieving state-of-the-art success rates and improved generalization to unseen goals. Our results suggest that quasimetric guidance is a viable alternative to RTG, opening new directions for learning structured distance representations in offline RL.

Degree

thesis:*
Name thesis:degree_name
Master of Science (MSc)
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Ontario Institute of Technology
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Goyani, Madhav
Advisors dc:contributor.advisor
  • Ebrahimi, Mehran
  • Davoudi, Kourosh

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10155/1960
OAI identifier oai:identifier
oai:ontariotechu.scholaris.ca:10155/1960

Chain of custody

source
Harvested from
Ontario Institute of Technology
Base URL
ontariotechu.scholaris.ca/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Goyani, Madhav. Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance. University of Ontario Institute of Technology, 2025. https://hdl.handle.net/10155/1960