Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 26 for “"Visual Question Answering"”.

  1. Visual question answering using external knowledge

    Accurately answering a question about a given image requires combining observations with general knowledge. While this is effortless for humans, reasoning with general knowledge remains an algorithmic challenge. To advance research in this direction, a novel `fact-based' visual question answering

    uiuc Repository record for Visual question answering using external knowledge (opens in a new tab)

  2. Role of Premises in Visual Question Answering

    … work, we make a simple but important observation questions about images often contain premises -- objects and relationships implied by the question -- and that reasoning about premises can help Visual Question Answering (VQA) models respond more intelligently to irrelevant or previously unseen …

    vt Repository record for Role of Premises in Visual Question Answering (opens in a new tab)

  3. Visual Question Answering in the Medical Domain

    … Thus, it becomes crucial to have a reliable Visual Question Answering (VQA) system which can provide a "second opinion" on medical cases. However, most of the VQA systems that work today cater to real-world problems and are not specifically tailored for handling medical images. Moreover, the …

    vt Repository record for Visual Question Answering in the Medical Domain (opens in a new tab)

  4. Fact-based visual question answering using knowledge graph embeddings

    … through vision and language. Fact-based Visual Question Answering (FVQA), a challenging variant of VQA, requires a QA-system to mimic this human ability. It must include facts from a diverse knowledge graph (KG) in its reasoning process to produce an answer. Large KGs, especially …

    uiuc Repository record for Fact-based visual question answering using knowledge graph embeddings (opens in a new tab)

  5. Augmenting Multi-modal Question Answering Systems with Retrieval Methods

    … generation (RAG) into multi-modal question answering (QA) systems as a solution to these challenges. By leveraging external knowledge sources, RAG enhances model accuracy and access to domain-specific information. The research unfolds in the following order: Firstly, to efficiently …

    cambridge Repository record for Augmenting Multi-modal Question Answering Systems with Retrieval Methods (opens in a new tab)

  6. Unifying cross-modal concepts in vision and language

    … amounts of data external to the target task. In visual question answering, models tend to rely on contextual cues or learned priors instead of actually recognizing and linking concepts across modalities. Consequently, when a concept appears in a new context, models often fail to adapt. We learn …

    uiuc Repository record for Unifying cross-modal concepts in vision and language (opens in a new tab)

  7. Evaluating visually grounded language capabilities using microworlds

    … as application-focused comparative benchmarks. Visual question answering is an example of a modern holistic understanding task, unifying a range of abilities around visually grounded language understanding in a single problem statement. It has also been an early example for which some of the …

    cambridge Repository record for Evaluating visually grounded language capabilities using microworlds (opens in a new tab)

  8. Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks

    … of commonsense knowledge while performing visual or textual tasks. Inspired by that ability, we approach commonsense learning as leveraging perspectives from multiple modalities for images and text in the context of vision and language tasks. Given a target task (e.g., textual reasoning, …

    vt Repository record for Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks (opens in a new tab)

  9. Data Augmentation with Seq2Seq Models

    … issue that complicates the training process of question answering systems: syntactically diverse but semantically equivalent sentences can have significant disparities in predicted output probabilities. We propose a method for generating an augmented paraphrase corpus for the visual question

    vt Repository record for Data Augmentation with Seq2Seq Models (opens in a new tab)

  10. Language Modeling from Visually Grounded Speech

    … fully leveraging multimodal inputs, particularly visual context, remains underexplored. This thesis addresses this gap by developing novel language modeling techniques directly from visually grounded speech. We first introduce the Audio-Visual Neural Syntax Learner (AV-NSL), an unsupervised parser …

    mit Repository record for Language Modeling from Visually Grounded Speech (opens in a new tab)

  11. Deep Attentional Modulation for Zero-shot Learning in Object Recognition

    … human ability to recognize seemingly infinite visual concepts using the same visual pathway. Even more impressive, humans have the ability to recognize objects from just a description (zero-shot) or a few examples (few-shot). Traditionally, artificial neural networks have struggled at …

    mit Repository record for Deep Attentional Modulation for Zero-shot Learning in Object Recognition (opens in a new tab)

  12. The Art of Deep Connection - Towards Natural and Pragmatic Conversational Agent Interactions

    … Although we now have agents that can answer questions asked for images, they are prone to failure from confusing input, and cannot ask clarification questions, in turn, to extract the desired information from humans. Hence, as a first step, we direct our efforts towards making Visual Question

    vt Repository record for The Art of Deep Connection - Towards Natural and Pragmatic Conversational Agent Interactions (opens in a new tab)

  13. Connecting vision and language via image retrieval and captioning

    … captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image retrieval is studied. In this problem, the objective is to pick the most related …

    manitoba Repository record for Connecting vision and language via image retrieval and captioning (opens in a new tab)

  14. Representations from vision and language

    … of concepts that manifest themselves visually, linguistically or through other senses. Furthermore concepts do not exist in isolation but are related to each other. In this work, we show how to build representations of concepts from visual and textual data, link visual manifestations …

    uiuc Repository record for Representations from vision and language (opens in a new tab)

  15. Advancing Chart Question Answering with Robust Chart Component Recognition

    … of key components, while the chart question answering (ChartQA) task integrates visual and textual information, facilitating accurate responses to queries based on the chart's content. To approach ChartQA, this research focuses on two main aspects. Firstly, we introduce ChartFormer, …

    vt Repository record for Advancing Chart Question Answering with Robust Chart Component Recognition (opens in a new tab)

  16. Towards Interpretable Vision Systems

    … tasks as examples, words semantic matching and Visual Question Answering (VQA). In VQA, we take binary questions on abstract scenes as the first stage, then we extend to all question types on real images. In both cases, we take attention as an important intermediate output. By explicitly forcing …

    vt Repository record for Towards Interpretable Vision Systems (opens in a new tab)

  17. Going Deeper with Images and Natural Language

    … is able to perceive and understand the complex visual environment around us. More ambitiously, it should be able to interact with us about its surroundings in natural languages. Thanks to the progress made in deep learning, we've seen huge breakthroughs towards this goal over the last few years. …

    vt Repository record for Going Deeper with Images and Natural Language (opens in a new tab)

  18. Transparent Analysis of Multi-Modal Embeddings

    … focuses on downstream tasks, involving direct visual input, such as Visual Question Answering. Fewer papers have exploited visual information for meaning representations when the evaluation tasks involve no direct visual input, such as semantic similarity. When such research has been …

    cambridge Repository record for Transparent Analysis of Multi-Modal Embeddings (opens in a new tab)

  19. Vision and Language: Information Integration and Transformation

    … multi-modalities, such as image-text retrieval, visual question answering, visual grounding, and image captioning. In particular, we explore and develop techniques combining image and text information to improve transformation between modalities. First, we exploit the multilingual image …

    rice Repository record for Vision and Language: Information Integration and Transformation (opens in a new tab)

  20. Methods for generating visual programs with optimizable vision models

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Methods for generating visual programs with optimizable vision models (opens in a new tab)

Page 1 of 2