Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 50 for “"Vision-language Models"”.

  1. Reasoning, scaling, generating with vision-language models

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Reasoning, scaling, generating with vision-language models (opens in a new tab)

  2. Model Control through Lightweight Activation Steering for Vision Language Models

    … a lightweight steering module designed to guide Vision- Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the …

    vt Repository record for Model Control through Lightweight Activation Steering for Vision Language Models (opens in a new tab)

  3. Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models

    … and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast …

    uoit Repository record for Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models (opens in a new tab)

  4. VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications

    Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly …

    mit Repository record for VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications (opens in a new tab)

  5. General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models

    … require AR system experts for use, require CAD models of real-world objects, or only function for limited types of tasks or environments. We propose a general-purpose AR task guidance approach and proof-of-concept system to generate guidance for tasks defined by natural language. Our approach …

    vt Repository record for General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models (opens in a new tab)

  6. Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing

    This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and …

    stellenbosch Repository record for Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing (opens in a new tab)

  7. Advanced Methods for Remote Sensing Image Captioning

    … captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …

    trento Repository record for Advanced Methods for Remote Sensing Image Captioning (opens in a new tab)

  8. SMALL AND LARGE PERCEPTION MODELS FOR ROBOTIC NAVIGATION

    Computer vision is fundamental to advancing robotic navigation and autonomous driving systems, enabling machines to interpret visual data critical for interaction with complex and diverse real-world environments. We address many critical challenges in these domains by developing advanced vision

    maryland Repository record for SMALL AND LARGE PERCEPTION MODELS FOR ROBOTIC NAVIGATION (opens in a new tab)

  9. Deep Learning Multimodal Extraction of Reaction Data

    … chemistry information extraction potential of Vision Language Models (VLM), which allow powerful large language models to leverage visual understanding. Our findings indicate that VLMs still require additional work in order to meet the performance of our bespoke models.

    mit Repository record for Deep Learning Multimodal Extraction of Reaction Data (opens in a new tab)

  10. Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps

    … proposes a methodology for implementing Large Language Model (LLM), and Vision Language Model (VLM) to enable delivery robots to identify the final delivery target and navigate the complex terrain from the curb to the front door. The proposed solution aims to enhance the autonomy and safety of …

    mit Repository record for Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps (opens in a new tab)

  11. Developing Domain-Specific Generative Methods

    … shifts to building on large-scale foundation models and powerful architectures, careful thought has to go into adapting these models to new domains and tasks. The work in this thesis demonstrates novel approaches to adapting large-scale generative models and architectures to specific …

    mit Repository record for Developing Domain-Specific Generative Methods (opens in a new tab)

  12. CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING

    … exhaustion, sustaining the improvement of large language models (LLMs) and large vision--language models (LVLMs) demands a paradigm shift. This thesis proposes automatic scaling: a closed-loop framework in which models autonomously improve through their own computation via three layers. …

    nus Repository record for CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING (opens in a new tab)

  13. Enhancing 3D Scene Graph Generation with Multimodal Embeddings

    … for scene understanding in robotics and computer vision. Current approaches for automated zero-shot 3D Scene Graph generation rely on spatial ontologies that relate objects with the semantic locations they are found in (e.g., a fork is found in a kitchen). While conferring impressive zero-shot …

    mit Repository record for Enhancing 3D Scene Graph Generation with Multimodal Embeddings (opens in a new tab)

  14. Visual Element Property Graphs for Bridging the Symbol Description-Recognition Gap

    … components and semantic relationships with a language model-powered natural language interface is developed. This method explicitly models relationships between visual elements and interpretations, differing from end-to-end vision-language models. Evaluations, using automated metrics and human …

    york Repository record for Visual Element Property Graphs for Bridging the Symbol Description-Recognition Gap (opens in a new tab)

  15. EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS

    … high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in …

    nus Repository record for EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS (opens in a new tab)

  16. AI-Based Framework for Identifying Wood Burning Appliances through Chimney Recognition

    … this limitation, we proposed a novel computer vision framework that leverages visible chimney features as proxies for wood stove usage. Using data from drones, vehicle-based videos across urban and rural areas, we trained object detection models (YOLOv11) and integrated them with Vision

    cornell Repository record for AI-Based Framework for Identifying Wood Burning Appliances through Chimney Recognition (opens in a new tab)

  17. A Taxonomic Framework for the Classification of Wildfire Social Media Data in Canada

    … systems during wildfire events. It evaluates vision-language models, deep learning, and traditional classifiers on this dataset, finding that custom deep learning models significantly outperform other classifiers, with the best achieving an f1-score of 84.48±0.69\%. The model’s capability to …

    carleton Repository record for A Taxonomic Framework for the Classification of Wildfire Social Media Data in Canada (opens in a new tab)

  18. Using AI to Improve Price Transparency in Real Estate Valuation

    … attributes to enhance traditional Hedonic models. By incorporating Vision Language Models (VLMs) and generative AI, the research evaluates the potential of these technologies to assess non-standard variables like aesthetic appeal, condition and cohesiveness of interior and exterior property …

    mit Repository record for Using AI to Improve Price Transparency in Real Estate Valuation (opens in a new tab)

  19. Model-based Planning for Efficient Task Execution

    … about both visual observations and high-level language instructions. However, they plan in a high-dimensional latent space, opaque to human collaborators. Hence, it is difficult for humans to understand the agent’s decision-making process. This lack of interpretability hinders effective …

    mit Repository record for Model-based Planning for Efficient Task Execution (opens in a new tab)

  20. Generalizable Robot Manipulation through Unified Perception, Policy Learning, and Planning

    … estimates affordances with learned perception models with task-and-motion-planning (TAMP) for object rearrangement in unstructured scenes, 2) learning generative diffusion models of robot skills, which can be composed to solve unseen combination of environmental constraints through …

    mit Repository record for Generalizable Robot Manipulation through Unified Perception, Policy Learning, and Planning (opens in a new tab)

Page 1 of 3