Back to results

City University of New York - City College

Cognitive Map Generation for Vision and Language Navigation

Abstract

dc:description.abstract

<p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems often struggle with effectively aligning language instructions and visual scene understanding, limiting their adaptability and navigation efficiency.</p> <p>This thesis proposes a novel Cognitive Map-based framework that addresses these challenges by transforming natural language navigation instructions into structured graph representations. The Cognitive Map consists of nodes representing waypoints, landmarks, decision points, and edges encoding spatial relationships and navigational actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.</p> <p>The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming algorithm is used to refine alignment for II segments with ambiguity or mismatch, leveraging semantic similarity computed via SBERT. This two-stage process ensures temporal coherence and semantic grounding between text-based instructions and real-world visual environments.</p> <p>Experiments conducted on the Touchdown and Map2Seq datasets demonstrate improved navigation accuracy, robust scene-language alignment, and increased adaptability compared to traditional VLN methods. The proposed Cognitive Map framework bridges the gap between visual and linguistic modalities, offering a scalable solution for real-world applications such as urban navigation, search and rescue, and augmented reality guidance systems.</p>

Degree

thesis:*
Name thesis:degree_name
Master of Science (M.S.)
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Year dc:date.available
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sandoval Mesa, Alexander
Contributors dc:contributor
  • Shigang Li
  • Zhigang Zhu

Subjects

dc:subject × 13

Identifiers

dc:identifier.*
Repository record dc:identifier
https://academicworks.cuny.edu/cc_etds_theses/1274
OAI identifier oai:identifier
oai:academicworks.cuny.edu:cc_etds_theses-2324

Chain of custody

source
Harvested from
City University of New York - City College
Base URL
academicworks.cuny.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Sandoval Mesa, Alexander. Cognitive Map Generation for Vision and Language Navigation. Thesis thesis, 2025. https://academicworks.cuny.edu/cc_etds_theses/1274