Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 1 of 1 for “"Multimodal Alignment"”.

  1. Cognitive Map Generation for Vision and Language Navigation

    … Google Street View are processed using Large Multimodal Models (LMMs) to generate comprehensive visual descriptions of the environment.</p> <p>The system first applies incremental alignment based on instruction order and heading direction to integrate these modalities. A dynamic programming …

    cuny Repository record for Cognitive Map Generation for Vision and Language Navigation (opens in a new tab)