Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 12 of 12 for “"Vision Language Model"”.

  1. A Submodular Approach to Find Interpretable Directions in Text-to-Image Models

    Text-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of …

    vt Repository record for A Submodular Approach to Find Interpretable Directions in Text-to-Image Models (opens in a new tab)

  2. Self-Supervised ECG Learning for Multimodal Clinical Tasks

    … with chest X-rays and EHR text using a visionlanguage model backbone, enabling end-to-end multimodal inference. Our results show that incorporating ECG signals meaningfully improves diagnostic performance, highlighting the value of multitask time series pretraining and modular fusion …

    mit Repository record for Self-Supervised ECG Learning for Multimodal Clinical Tasks (opens in a new tab)

  3. Neural Feature Fields for Language-Guided Robot Manipulation

    Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic …

    mit Repository record for Neural Feature Fields for Language-Guided Robot Manipulation (opens in a new tab)

  4. DYNAMIC NEURAL NETWORKS FOR EFFICIENT VISION MODEL INFERENCE

    … and 1.7× speedup with improved quality. Vision-Language Model (VLM) Acceleration: By implementing a dynamic small-large model cooperation mechanism, we achieve up to 3× faster inference while maintaining competitive performance. In conclusion, this research provides a comprehensive …

    nus Repository record for DYNAMIC NEURAL NETWORKS FOR EFFICIENT VISION MODEL INFERENCE (opens in a new tab)

  5. Advanced Methods for Remote Sensing Image Captioning

    … captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …

    trento Repository record for Advanced Methods for Remote Sensing Image Captioning (opens in a new tab)

  6. Domain-Independent Mode Estimation for Human-Robot Collaboration

    … noise RGB-D data. It resolves ambiguity using Vision-Language Model (VLM)-guided semantic arbitration and demonstrates robustness and adaptability in unstructured environments. This work establishes qualitative spatial filtering with A*BC as a generalizable and efficient solution for semantic …

    mit Repository record for Domain-Independent Mode Estimation for Human-Robot Collaboration (opens in a new tab)

  7. Enriching models of natural language with auxiliary data

    The highest-performing natural language processing models generally solve language tasks by deriving statistical regularities of sequences of arbitrary tokens supplied as training data. Humans have a much richer notion of language, however. For one thing, they understand that language refers to …

    mit Repository record for Enriching models of natural language with auxiliary data (opens in a new tab)

  8. Visually Accurate Database-Enabled Reconstructions of Scenes (VADERS)

    … segments and rescaled per-component using vision-language model predictions to match the reference object better. Finally the asset is retextured based on the image mask of the reference object in the input image. Evaluation on six diverse scenes—both photographs and artworks—shows the …

    mit Repository record for Visually Accurate Database-Enabled Reconstructions of Scenes (VADERS) (opens in a new tab)

  9. Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing

    This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and …

    stellenbosch Repository record for Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing (opens in a new tab)

  10. Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps

    … proposes a methodology for implementing Large Language Model (LLM), and Vision Language Model (VLM) to enable delivery robots to identify the final delivery target and navigate the complex terrain from the curb to the front door. The proposed solution aims to enhance the autonomy and safety of …

    mit Repository record for Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps (opens in a new tab)

  11. Advanced Radar Sounder Data Analysis Methods under Limited labeled Data Constraints

    … that can operate effectively under weak supervision and limited labeled data, while also generalizing across different environments. This thesis addresses these challenges through a progression of frameworks that advance from classical deep learning to foundation models, with an emphasis on …

    trento Repository record for Advanced Radar Sounder Data Analysis Methods under Limited labeled Data Constraints (opens in a new tab)

  12. Improving medical report generation and evaluation through prompt engineering

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01

    uiuc Repository record for Improving medical report generation and evaluation through prompt engineering (opens in a new tab)