Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 83 for “"Vision-Language"”.

  1. Reasoning, scaling, generating with vision-language models

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Reasoning, scaling, generating with vision-language models (opens in a new tab)

  2. Chart Question Answering with an Universal Vision-Language Pretraining Approach

    … chart-based data analysis using natural language, several downstream tasks have been introduced recently including chart question answering. However, existing methods for these tasks often rely on pretraining on language or vision-language tasks, neglecting the explicit modeling of chart …

    york Repository record for Chart Question Answering with an Universal Vision-Language Pretraining Approach (opens in a new tab)

  3. Model Control through Lightweight Activation Steering for Vision Language Models

    … a lightweight steering module designed to guide Vision- Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the …

    vt Repository record for Model Control through Lightweight Activation Steering for Vision Language Models (opens in a new tab)

  4. VoxelPrompt: A Vision-Language Agent for Grounded Medical Image Analysis

    We present VoxelPrompt, an agent-driven vision-language framework that tackles diverse radiological tasks through joint modeling of natural language, image volumes, and analytical metrics. VoxelPrompt is multi-modal and versatile, leveraging the flexibility of language interaction while providing …

    mit Repository record for VoxelPrompt: A Vision-Language Agent for Grounded Medical Image Analysis (opens in a new tab)

  5. Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models

    Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific …

    uoit Repository record for Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models (opens in a new tab)

  6. VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications

    Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly …

    mit Repository record for VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications (opens in a new tab)

  7. General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models

    … generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance …

    vt Repository record for General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models (opens in a new tab)

  8. Minimalist Approach to End-to-End Vision Language Navigation with Multi-Modal Foundation Model Features

    Recent vision-language navigation (VLN) approaches leverage large models, prompt engineering, and/or explicit reasoning for instruction interpretation and agent guidance. We introduce MiniNav, a minimalist framework employing frozen vision-language foundation models as patch-wise feature …

    mit Repository record for Minimalist Approach to End-to-End Vision Language Navigation with Multi-Modal Foundation Model Features (opens in a new tab)

  9. Quake-OVUDA: Component-Level AI-Based Post-Earthquake Building Inspections with VisionLanguage Guided Unsupervised Domain Adaptation

    … by leveraging the vast context provided by vision language models (VLMs). The proposed framework, termed Quake-OVUDA, and its additional supplementation with VLM expertise, outperform direct segmentation, open-vocabulary segmentation, and baseline UDA methods for building component and …

    houston Repository record for Quake-OVUDA: Component-Level AI-Based Post-Earthquake Building Inspections with Vision–Language Guided Unsupervised Domain Adaptation (opens in a new tab)

  10. Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing

    This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and …

    stellenbosch Repository record for Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing (opens in a new tab)

  11. Advanced Methods for Remote Sensing Image Captioning

    … captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …

    trento Repository record for Advanced Methods for Remote Sensing Image Captioning (opens in a new tab)

  12. Model-based Planning for Efficient Task Execution

    … about both visual observations and high-level language instructions. However, they plan in a high-dimensional latent space, opaque to human collaborators. Hence, it is difficult for humans to understand the agent’s decision-making process. This lack of interpretability hinders effective …

    mit Repository record for Model-based Planning for Efficient Task Execution (opens in a new tab)

  13. Connecting Deep Learning Models to the Human Brain

    … models, particularly models that integrate vision and language with human brain processing. These models have shown remarkable advancements in tasks such as object recognition, scene classification, and language processing, achieving near-human accuracy in some cases. This raises intriguing …

    mit Repository record for Connecting Deep Learning Models to the Human Brain (opens in a new tab)

  14. Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps

    … proposes a methodology for implementing Large Language Model (LLM), and Vision Language Model (VLM) to enable delivery robots to identify the final delivery target and navigate the complex terrain from the curb to the front door. The proposed solution aims to enhance the autonomy and safety of …

    mit Repository record for Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps (opens in a new tab)

  15. SMALL AND LARGE PERCEPTION MODELS FOR ROBOTIC NAVIGATION

    Computer vision is fundamental to advancing robotic navigation and autonomous driving systems, enabling machines to interpret visual data critical for interaction with complex and diverse real-world environments. We address many critical challenges in these domains by developing advanced vision

    maryland Repository record for SMALL AND LARGE PERCEPTION MODELS FOR ROBOTIC NAVIGATION (opens in a new tab)

  16. Instruction Mining from Images: Constructing a Synthetic Dataset for Multimodal Learning

    … εξέλιξη των μεγάλων γλωσσικών μοντέλων (Large Language Models – LLMs) και των πολυτροπικών μοντέλων όρασης–γλώσσας (Multimodal Large Language Models – MLLMs) έχει οδηγήσει στην ανάπτυξη συστημάτων ικανών να συνδυάζουν οπτική και γλωσσική πληροφορία για την εκτέλεση σύνθετων εργασιών …

    athens Repository record for Instruction Mining from Images: Constructing a Synthetic Dataset for Multimodal Learning (opens in a new tab)

  17. Grounded Language Learning with Foundation Models

    The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array …

    cambridge Repository record for Grounded Language Learning with Foundation Models (opens in a new tab)

  18. Language-Guided Video Understanding with Foundation Models

    … remains limited by assumptions about supervision, training data availability, and offline access to complete video sequences. These constraints are particularly restrictive in settings such as surveillance and procedural assistance, where data is scarce, privacy-sensitive, and decisions …

    trento Repository record for Language-Guided Video Understanding with Foundation Models (opens in a new tab)

Page 1 of 5