Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 7 of 7 for “"Large Vision-Language Models"”.

  1. CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING

    … exhaustion, sustaining the improvement of large language models (LLMs) and large vision--language models (LVLMs) demands a paradigm shift. This thesis proposes automatic scaling: a closed-loop framework in which models autonomously improve through their own computation via three layers. …

    nus Repository record for CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING (opens in a new tab)

  2. VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications

    Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions. However, it remains unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly …

    mit Repository record for VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications (opens in a new tab)

  3. EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS

    … high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in …

    nus Repository record for EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS (opens in a new tab)

  4. Generalizable Robot Manipulation through Unified Perception, Policy Learning, and Planning

    … research. While the availability of data and large-scale training has brought exciting progress in robotics manipulation, current methods often struggle with generalizing to unseen, unstructured environments and solving long-horizon tasks. In this thesis, I will present my work in robot …

    mit Repository record for Generalizable Robot Manipulation through Unified Perception, Policy Learning, and Planning (opens in a new tab)

  5. Advanced Methods for Remote Sensing Image Captioning

    … captioning (IC) involves generating natural language descriptions for images, enabling machines to communicate their perception through language. This approach provides a flexible framework to convey diverse semantics. Despite significant advancements, IC systems face challenges in …

    trento Repository record for Advanced Methods for Remote Sensing Image Captioning (opens in a new tab)

  6. Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation

    … laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal …

    vt Repository record for Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation (opens in a new tab)

  7. Scalable foundation models

    … opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand …

    uiuc Repository record for Scalable foundation models (opens in a new tab)