Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 7 of 7 for “"Visual question answering (VQA)"”.

  1. Role of Premises in Visual Question Answering

    … work, we make a simple but important observation questions about images often contain premises -- objects and relationships implied by the question -- and that reasoning about premises can help Visual Question Answering (VQA) models respond more intelligently to irrelevant or previously unseen …

    vt Repository record for Role of Premises in Visual Question Answering (opens in a new tab)

  2. Visual Question Answering in the Medical Domain

    … Thus, it becomes crucial to have a reliable Visual Question Answering (VQA) system which can provide a "second opinion" on medical cases. However, most of the VQA systems that work today cater to real-world problems and are not specifically tailored for handling medical images. Moreover, the …

    vt Repository record for Visual Question Answering in the Medical Domain (opens in a new tab)

  3. Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks

    … of commonsense knowledge while performing visual or textual tasks. Inspired by that ability, we approach commonsense learning as leveraging perspectives from multiple modalities for images and text in the context of vision and language tasks. Given a target task (e.g., textual reasoning, …

    vt Repository record for Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks (opens in a new tab)

  4. Towards Interpretable Vision Systems

    … tasks as examples, words semantic matching and Visual Question Answering (VQA). In VQA, we take binary questions on abstract scenes as the first stage, then we extend to all question types on real images. In both cases, we take attention as an important intermediate output. By explicitly forcing …

    vt Repository record for Towards Interpretable Vision Systems (opens in a new tab)

  5. Exploiting multimodality and structure in world representations

    … concern various aspects of such systems---visual reasoning, language representations, causal mechanisms, robustness to out-of-distribution inputs, to name only a few. In particular, multimodal learning and language grounding are vital to achieving a strong understanding of the real world. …

    cambridge Repository record for Exploiting multimodality and structure in world representations (opens in a new tab)

  6. Learning visual tasks with selective attention

    … resulting in significant gains in several visual prediction tasks. We will demonstrate both directly and indirectly supervised models for selecting image regions and show how they can improve performance over baselines by means of focusing on the right areas.

    uiuc Repository record for Learning visual tasks with selective attention (opens in a new tab)

  7. Unifying cross-modal concepts in vision and language

    … amounts of data external to the target task. In visual question answering, models tend to rely on contextual cues or learned priors instead of actually recognizing and linking concepts across modalities. Consequently, when a concept appears in a new context, models often fail to adapt. We learn …

    uiuc Repository record for Unifying cross-modal concepts in vision and language (opens in a new tab)