Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 20 for “"Image Captioning"”.

  1. Automatic Image Captioning with Style

    … vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …

    aus-cath Repository record for Automatic Image Captioning with Style (opens in a new tab)

  2. Automatic Image Captioning with Style

    … vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …

    anu Repository record for Automatic Image Captioning with Style (opens in a new tab)

  3. Image captioning using compositional sentiments

    … a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these …

    uiuc Repository record for Image captioning using compositional sentiments (opens in a new tab)

  4. IMAGE CAPTIONING FOR REMOTE SENSING IMAGE ANALYSIS

    Image Captioning (IC) aims to generate a coherent and comprehensive textual description that summarizes the complex content of an image. It is a combination of computer vision and natural language processing techniques to encode the visual features of an image and translate them into a sentence. In …

    trento Repository record for IMAGE CAPTIONING FOR REMOTE SENSING IMAGE ANALYSIS (opens in a new tab)

  5. Advanced Methods for Remote Sensing Image Captioning

    Someone once said that ”an image is worth a thousand words.” This captures well the amount of semantic information hidden inside these matrices of pixels. Looking at an image instinctively induces us to form hypotheses about which objects are inside, their state, dislocation, etc. We thus create, …

    trento Repository record for Advanced Methods for Remote Sensing Image Captioning (opens in a new tab)

  6. Exposing and correcting the gender bias in image captioning datasets and models

    The task of image captioning implicitly involves gender identification. However, due to the gender bias in data, gender identification by an image captioning model suffers. Also, due to the word-by-word prediction, the gender-activity bias in the data tends to influence the other words in the …

    uiuc Repository record for Exposing and correcting the gender bias in image captioning datasets and models (opens in a new tab)

  7. Connecting vision and language via image retrieval and captioning

    … jointly understanding vision and language, e.g. image/video captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image retrieval is studied. In this …

    manitoba Repository record for Connecting vision and language via image retrieval and captioning (opens in a new tab)

  8. Exploring the Internal Statistics: Single Image Super-Resolution, Completion and Captioning

    <p>Image enhancement has drawn increasingly attention in improving image quality or interpretability. It aims to modify images to achieve a better perception for human visual system or a more suitable representation for further analysis in a variety of applications such as medical imaging, remote …

    cuny-grad Repository record for Exploring the Internal Statistics: Single Image Super-Resolution, Completion and Captioning (opens in a new tab)

  9. Vision and Language: Information Integration and Transformation

    Images and text are large-scale data sources for deep learning models to mimic how humans perceive and understand multimodal information. Combining images and text can construct better feature representations by involving complementary information from different sources. For example, text …

    rice Repository record for Vision and Language: Information Integration and Transformation (opens in a new tab)

  10. Learning multiple solutions to computer vision problems

    … ones are the progress made on the problems of image classification [8, 9, 10], image segmentation [11, 12, 13, 14, 15], object detection [16, 17, 12, 18] and vision language tasks, e.g. image captioning [19, 20, 21], visual question answering [22, 23, 24, 25, 26] etc. Convolutional neural …

    uiuc Repository record for Learning multiple solutions to computer vision problems (opens in a new tab)

  11. Multimodal machine translation

    … is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this …

    uiuc Repository record for Multimodal machine translation (opens in a new tab)

  12. Learning joint latent representations for images and language

    … Learning the joint latent representations for images and language is vital to solving many image-text tasks, including image-sentence retrieval, visual grounding, and image captioning, etc. In this thesis, we first propose two-branch neural networks for learning the similarity between these two …

    uiuc Repository record for Learning joint latent representations for images and language (opens in a new tab)

  13. VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications

    … unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly satellite and aerial images less common in VLM training data. In this work, we propose VLEO-Bench, a comprehensive evaluation framework to quantify the progress of VLMs toward …

    mit Repository record for VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications (opens in a new tab)

  14. Topic-Based Video Classification and Retrieval Using Machine Learning

    … primarily concentrate on object detection, image classification, and image captioning. However, very little work has been shown in DL-based video-content analysis and retrieval. Due to the complex nature of time relevant information in a sequence of video frames, understanding video contents …

    umkc Repository record for Topic-Based Video Classification and Retrieval Using Machine Learning (opens in a new tab)

  15. Evaluating visually grounded language capabilities using microworlds

    … a new approach to assess generative tasks like image captioning.

    cambridge Repository record for Evaluating visually grounded language capabilities using microworlds (opens in a new tab)

  16. Evaluating Natural Language Generation Tasks for Grammaticality, Faithfulness and Diversity

    … generation tasks are explored: synthetic image captioning, football highlight generation from match statistics, and topic-shift dialogue generation. These tasks are deliberately chosen to cover a diverse range of generation scenarios. Each task provides unique grounding information and …

    cambridge Repository record for Evaluating Natural Language Generation Tasks for Grammaticality, Faithfulness and Diversity (opens in a new tab)

  17. Going Deeper with Images and Natural Language

    … in which machines now can categorize images into multiple classes, and detect various objects within an image, with an ability that is competitive with or even surpasses that of humans. Meanwhile, we also have witnessed similar strides in natural language processing (NLP). It is quite …

    vt Repository record for Going Deeper with Images and Natural Language (opens in a new tab)

  18. Evaluating interactions and decision-making by annotators and users in computer vision systems

    … helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the …

    temple Repository record for Evaluating interactions and decision-making by annotators and users in computer vision systems (opens in a new tab)

  19. Learning to map between domains

    … something seen before. Similarly, in the medical image field and radiological science, tens of thousands of medical images (MRI, CT, etc) of patients are taken. These medical images need to be studied and interpreted. In this dissertation, we investigate a number of data-driven approaches for …

    uiuc Repository record for Learning to map between domains (opens in a new tab)

  20. Natural language image description: data, models, and evaluation

    Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: …

    uiuc Repository record for Natural language image description: data, models, and evaluation (opens in a new tab)