{"id":{"repo_id":"cuny","oai_identifier":"oai:academicworks.cuny.edu:cc_etds_theses-2324"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny/oai:academicworks.cuny.edu:cc_etds_theses-2324","repository":{"repo_id":"cuny","name":"City University of New York - City College","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Cognitive Map Generation for Vision and Language Navigation","abstract":"<p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems often struggle with effectively aligning language instructions and visual scene understanding, limiting their adaptability and navigation efficiency.</p> <p>This thesis proposes a novel Cognitive Map-based framework that addresses these challenges by transforming natural language navigation instructions into structured graph representations. The Cognitive Map consists of nodes representing waypoints, landmarks, decision points, and edges encoding spatial relationships and navigational actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.</p> <p>The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming algorithm is used to refine alignment for II segments with ambiguity or mismatch, leveraging semantic similarity computed via SBERT. This two-stage process ensures temporal coherence and semantic grounding between text-based instructions and real-world visual environments.</p> <p>Experiments conducted on the Touchdown and Map2Seq datasets demonstrate improved navigation accuracy, robust scene-language alignment, and increased adaptability compared to traditional VLN methods. The proposed Cognitive Map framework bridges the gap between visual and linguistic modalities, offering a scalable solution for real-world applications such as urban navigation, search and rescue, and augmented reality guidance systems.</p>","abstract_html":"&lt;p&gt;Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems often struggle with effectively aligning language instructions and visual scene understanding, limiting their adaptability and navigation efficiency.&lt;/p&gt; &lt;p&gt;This thesis proposes a novel Cognitive Map-based framework that addresses these challenges by transforming natural language navigation instructions into structured graph representations. The Cognitive Map consists of nodes representing waypoints, landmarks, decision points, and edges encoding spatial relationships and navigational actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.&lt;/p&gt; &lt;p&gt;The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming algorithm is used to refine alignment for II segments with ambiguity or mismatch, leveraging semantic similarity computed via SBERT. This two-stage process ensures temporal coherence and semantic grounding between text-based instructions and real-world visual environments.&lt;/p&gt; &lt;p&gt;Experiments conducted on the Touchdown and Map2Seq datasets demonstrate improved navigation accuracy, robust scene-language alignment, and increased adaptability compared to traditional VLN methods. The proposed Cognitive Map framework bridges the gap between visual and linguistic modalities, offering a scalable solution for real-world applications such as urban navigation, search and rescue, and augmented reality guidance systems.&lt;/p&gt;","abstract_has_math":false,"creators":["Sandoval Mesa, Alexander"],"institution":null,"degree_name":"Master of Science (M.S.)","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Shigang Li","Zhigang Zhu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-01-01T08:00:00Z","date_published":"2025-01-01T08:00:00Z","updated_at":"2026-07-24T01:58:13Z","subjects":["Cognitive Maps","Visual-Language Navigation","Multimodal Alignment","Semantic Matching","Dynamic Programming","Urban Navigation","Scene Interpretation Engineering","Artificial Intelligence and Robotics","Computational Engineering","Data Science","Geological Engineering","Other Computer Engineering","Transportation Engineering"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/cc_etds_theses/1274","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shigang Li","Zhigang Zhu"]},{"key":"dc:creator","label":"Author","values":["Sandoval Mesa, Alexander"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2026-05-20T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (M.S.)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cognitive Maps","Visual-Language Navigation","Multimodal Alignment","Semantic Matching","Dynamic Programming","Urban Navigation","Scene Interpretation Engineering","Artificial Intelligence and Robotics","Computational Engineering","Data Science","Geological Engineering","Other Computer Engineering","Transportation Engineering"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/cc_etds_theses/1274"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems often struggle with effectively aligning language instructions and visual scene understanding, limiting their adaptability and navigation efficiency.</p> <p>This thesis proposes a novel Cognitive Map-based framework that addresses these challenges by transforming natural language navigation instructions into structured graph representations. The Cognitive Map consists of nodes representing waypoints, landmarks, decision points, and edges encoding spatial relationships and navigational actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.</p> <p>The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming algorithm is used to refine alignment for II segments with ambiguity or mismatch, leveraging semantic similarity computed via SBERT. This two-stage process ensures temporal coherence and semantic grounding between text-based instructions and real-world visual environments.</p> <p>Experiments conducted on the Touchdown and Map2Seq datasets demonstrate improved navigation accuracy, robust scene-language alignment, and increased adaptability compared to traditional VLN methods. The proposed Cognitive Map framework bridges the gap between visual and linguistic modalities, offering a scalable solution for real-world applications such as urban navigation, search and rescue, and augmented reality guidance systems.</p>"]},{"key":"dc:title","label":"Title","values":["Cognitive Map Generation for Vision and Language Navigation"]}]}],"canonical_facts":{"dc:contributor":["Shigang Li","Zhigang Zhu"],"dc:creator":["Sandoval Mesa, Alexander"],"dc:date.available":["2026-05-20T07:00:00Z"],"dc:description.abstract":["<p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems often struggle with effectively aligning language instructions and visual scene understanding, limiting their adaptability and navigation efficiency.</p> <p>This thesis proposes a novel Cognitive Map-based framework that addresses these challenges by transforming natural language navigation instructions into structured graph representations. The Cognitive Map consists of nodes representing waypoints, landmarks, decision points, and edges encoding spatial relationships and navigational actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.</p> <p>The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming algorithm is used to refine alignment for II segments with ambiguity or mismatch, leveraging semantic similarity computed via SBERT. This two-stage process ensures temporal coherence and semantic grounding between text-based instructions and real-world visual environments.</p> <p>Experiments conducted on the Touchdown and Map2Seq datasets demonstrate improved navigation accuracy, robust scene-language alignment, and increased adaptability compared to traditional VLN methods. The proposed Cognitive Map framework bridges the gap between visual and linguistic modalities, offering a scalable solution for real-world applications such as urban navigation, search and rescue, and augmented reality guidance systems.</p>"],"dc:identifier":["https://academicworks.cuny.edu/cc_etds_theses/1274"],"dc:subject":["Cognitive Maps","Visual-Language Navigation","Multimodal Alignment","Semantic Matching","Dynamic Programming","Urban Navigation","Scene Interpretation Engineering","Artificial Intelligence and Robotics","Computational Engineering","Data Science","Geological Engineering","Other Computer Engineering","Transportation Engineering"],"dc:title":["Cognitive Map Generation for Vision and Language Navigation"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (M.S.)"]},"updated_at":"2026-07-24T01:58:13Z"}