Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 7 of 7 for “"Large Vision-Language Models"”.
-
CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING
… exhaustion, sustaining the improvement of large language models (LLMs) and large vision--language models (LVLMs) demands a paradigm shift. This thesis proposes automatic scaling: a closed-loop framework in which models autonomously improve through their own computation via three layers. …
-
VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications
Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly …
-
EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS
… high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in …
-
Generalizable Robot Manipulation through Unified Perception, Policy Learning, and Planning
… research. While the availability of data and large-scale training has brought exciting progress in robotics manipulation, current methods often struggle with generalizing to unseen, unstructured environments and solving long-horizon tasks. In this thesis, I will present my work in robot …
-
Advanced Methods for Remote Sensing Image Captioning
… captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …
-
Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation
… laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal …
-
Scalable foundation models
… opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand …