Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 83 for “"Vision-Language"”.
-
Reasoning, scaling, generating with vision-language models
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms
-
Chart Question Answering with an Universal Vision-Language Pretraining Approach
… chart-based data analysis using natural language, several downstream tasks have been introduced recently including chart question answering. However, existing methods for these tasks often rely on pretraining on language or vision-language tasks, neglecting the explicit modeling of chart …
-
Model Control through Lightweight Activation Steering for Vision Language Models
… a lightweight steering module designed to guide Vision- Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the …
-
VoxelPrompt: A Vision-Language Agent for Grounded Medical Image Analysis
We present VoxelPrompt, an agent-driven vision-language framework that tackles diverse radiological tasks through joint modeling of natural language, image volumes, and analytical metrics. VoxelPrompt is multi-modal and versatile, leveraging the flexibility of language interaction while providing …
-
Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models
Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific …
-
VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications
Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly …
-
General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models
… generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance …
-
Minimalist Approach to End-to-End Vision Language Navigation with Multi-Modal Foundation Model Features
Recent vision-language navigation (VLN) approaches leverage large models, prompt engineering, and/or explicit reasoning for instruction interpretation and agent guidance. We introduce MiniNav, a minimalist framework employing frozen vision-language foundation models as patch-wise feature …
-
Quake-OVUDA: Component-Level AI-Based Post-Earthquake Building Inspections with Vision–Language Guided Unsupervised Domain Adaptation
… by leveraging the vast context provided by vision language models (VLMs). The proposed framework, termed Quake-OVUDA, and its additional supplementation with VLM expertise, outperform direct segmentation, open-vocabulary segmentation, and baseline UDA methods for building component and …
-
Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video
Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01
-
Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing
This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and …
-
Advanced Methods for Remote Sensing Image Captioning
… captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …
-
Model-based Planning for Efficient Task Execution
… about both visual observations and high-level language instructions. However, they plan in a high-dimensional latent space, opaque to human collaborators. Hence, it is difficult for humans to understand the agent’s decision-making process. This lack of interpretability hinders effective …
-
Connecting Deep Learning Models to the Human Brain
… models, particularly models that integrate vision and language with human brain processing. These models have shown remarkable advancements in tasks such as object recognition, scene classification, and language processing, achieving near-human accuracy in some cases. This raises intriguing …
-
Last-Meter Delivery: Solving the Unattended Delivery Challenge from Streets to Doorsteps
… proposes a methodology for implementing Large Language Model (LLM), and Vision Language Model (VLM) to enable delivery robots to identify the final delivery target and navigate the complex terrain from the curb to the front door. The proposed solution aims to enhance the autonomy and safety of …
-
SMALL AND LARGE PERCEPTION MODELS FOR ROBOTIC NAVIGATION
Computer vision is fundamental to advancing robotic navigation and autonomous driving systems, enabling machines to interpret visual data critical for interaction with complex and diverse real-world environments. We address many critical challenges in these domains by developing advanced vision …
-
Instruction Mining from Images: Constructing a Synthetic Dataset for Multimodal Learning
… εξέλιξη των μεγάλων γλωσσικών μοντέλων (Large Language Models – LLMs) και των πολυτροπικών μοντέλων όρασης–γλώσσας (Multimodal Large Language Models – MLLMs) έχει οδηγήσει στην ανάπτυξη συστημάτων ικανών να συνδυάζουν οπτική και γλωσσική πληροφορία για την εκτέλεση σύνθετων εργασιών …
-
Grounded Language Learning with Foundation Models
The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array …
-
Language-Guided Video Understanding with Foundation Models
… remains limited by assumptions about supervision, training data availability, and offline access to complete video sequences. These constraints are particularly restrictive in settings such as surveillance and procedural assistance, where data is scarce, privacy-sensitive, and decisions …
Page 1 of 5