Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 8 of 8 for “"Image Caption"”.

  1. Textual entailment from image caption denotations

    … by building representations from grounded image captions. This allows us to use descriptions of the world to learn connections that would be difficult to identify in text-based corpora. In particular, we explore novel approaches to entailment that capture everyday world knowledge missing …

    uiuc Repository record for Textual entailment from image caption denotations (opens in a new tab)

  2. Automatic Image Captioning with Style

    … vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …

    aus-cath Repository record for Automatic Image Captioning with Style (opens in a new tab)

  3. Automatic Image Captioning with Style

    … vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …

    anu Repository record for Automatic Image Captioning with Style (opens in a new tab)

  4. Using the visual denotations of image captions for semantic inference

    … denotation of an utterance, which is the set of images that it describes. This notion borrows the abstract concept of a denotation of an utterance as the set of possible worlds in which the utterance is true from the logic-based approach, and instantiates possible worlds as images. In this …

    uiuc Repository record for Using the visual denotations of image captions for semantic inference (opens in a new tab)

  5. Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images

    … method for either training new language-aware image encoders or augmenting existing pretrained models with zero-shot visual recognition capabilities. However, existing works typically train on large datasets of image-text pairs and have been designed to perform downstream tasks involving only …

    mit Repository record for Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images (opens in a new tab)

  6. Connecting vision and language via image retrieval and captioning

    … jointly understanding vision and language, e.g. image/video captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image retrieval is studied. In this …

    manitoba Repository record for Connecting vision and language via image retrieval and captioning (opens in a new tab)

  7. Learning to map between domains

    … something seen before. Similarly, in the medical image field and radiological science, tens of thousands of medical images (MRI, CT, etc) of patients are taken. These medical images need to be studied and interpreted. In this dissertation, we investigate a number of data-driven approaches for …

    uiuc Repository record for Learning to map between domains (opens in a new tab)

  8. Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks

    … Humans have the remarkable ability to interpret images and text from different perspectives in multiple modalities, and to use large amounts of commonsense knowledge while performing visual or textual tasks. Inspired by that ability, we approach commonsense learning as leveraging perspectives …

    vt Repository record for Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks (opens in a new tab)