Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 7 of 7 for “"Large Multimodal Models"”.

  1. Learning without Labels - Reducing Supervision in Training, Inference, and Evaluation of Deep Neural Networks

    … fixed output vocabularies from Vision Language Models by formalizing the tasks of Vocabulary-free Image Classification and Vocabulary-free Semantic Segmentation and by introducing a family of efficient methods that adapt CLIP to the tasks. We also evaluate Large Multimodal Models under a similar …

    trento Repository record for Learning without Labels - Reducing Supervision in Training, Inference, and Evaluation of Deep Neural Networks (opens in a new tab)

  2. A Pedagogical Multimodal System for Mathematical Problem-Solving and Visual Reasoning

    … theories. This thesis explores and suggests how multimodal interaction between humans and AI helps humans engage with the system more naturally and effectively, leading to improved problem-solving in mathematical settings. Recent large multimodal models (LMMs) have the ability to facilitate …

    mit Repository record for A Pedagogical Multimodal System for Mathematical Problem-Solving and Visual Reasoning (opens in a new tab)

  3. Cognitive Map Generation for Vision and Language Navigation

    … actions. These maps are generated using Large Language Models (LLMs), specifically GPT-4o, to extract spatial information and construct detailed route topologies. Additionally, panoramic images sourced from Google Street View are processed using Large Multimodal Models (LMMs) to generate …

    cuny Repository record for Cognitive Map Generation for Vision and Language Navigation (opens in a new tab)

  4. Graph-based Vector Search Algorithms for Retrieval-Augmented AI Systems

    The recent advancement of large language models (LLMs) and large multimodal models (LMMs) greatly enhances the capabilities of AI systems such as recommendation systems and coding assistants, making them more practical for real-world deployment. However, these models cannot directly interact with …

    mit Repository record for Graph-based Vector Search Algorithms for Retrieval-Augmented AI Systems (opens in a new tab)

  5. Language-Guided Video Understanding with Foundation Models

    … decisions must be made online. Recent foundation models provide a new opportunity to rethink video understanding as a language-guided inference problem. Leveraging this shift, this thesis investigates how Vision-Language Models (VLMs) and Large Language Models (LLMs) can be used to relax key …

    trento Repository record for Language-Guided Video Understanding with Foundation Models (opens in a new tab)

  6. Exploring screen summarization with large language and multimodal models

    Mobile UIs are inherently multimodal. They can be represented by both visual representations, (screenshots) and structural metadata (view hierarchies). The image modality is rich and can contain information such as images, colors and positional information, while the view-hierarchies represent a …

    uiuc Repository record for Exploring screen summarization with large language and multimodal models (opens in a new tab)