Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 20 for “"Image Captioning"”.
-
Automatic Image Captioning with Style
… vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …
-
Automatic Image Captioning with Style
… vision and language. The problem of choice is image caption generation: automatically constructing natural language descriptions of image content. Previous research into image caption generation has focused on generating purely descriptive captions; I focus on generating visually relevant …
-
Image captioning using compositional sentiments
… a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these …
-
IMAGE CAPTIONING FOR REMOTE SENSING IMAGE ANALYSIS
Image Captioning (IC) aims to generate a coherent and comprehensive textual description that summarizes the complex content of an image. It is a combination of computer vision and natural language processing techniques to encode the visual features of an image and translate them into a sentence. In …
-
Advanced Methods for Remote Sensing Image Captioning
Someone once said that ”an image is worth a thousand words.” This captures well the amount of semantic information hidden inside these matrices of pixels. Looking at an image instinctively induces us to form hypotheses about which objects are inside, their state, dislocation, etc. We thus create, …
-
Exposing and correcting the gender bias in image captioning datasets and models
The task of image captioning implicitly involves gender identification. However, due to the gender bias in data, gender identification by an image captioning model suffers. Also, due to the word-by-word prediction, the gender-activity bias in the data tends to influence the other words in the …
-
Connecting vision and language via image retrieval and captioning
… jointly understanding vision and language, e.g. image/video captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image retrieval is studied. In this …
-
Exploring the Internal Statistics: Single Image Super-Resolution, Completion and Captioning
<p>Image enhancement has drawn increasingly attention in improving image quality or interpretability. It aims to modify images to achieve a better perception for human visual system or a more suitable representation for further analysis in a variety of applications such as medical imaging, remote …
-
Vision and Language: Information Integration and Transformation
Images and text are large-scale data sources for deep learning models to mimic how humans perceive and understand multimodal information. Combining images and text can construct better feature representations by involving complementary information from different sources. For example, text …
-
Learning multiple solutions to computer vision problems
… ones are the progress made on the problems of image classification [8, 9, 10], image segmentation [11, 12, 13, 14, 15], object detection [16, 17, 12, 18] and vision language tasks, e.g. image captioning [19, 20, 21], visual question answering [22, 23, 24, 25, 26] etc. Convolutional neural …
-
Multimodal machine translation
… is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this …
-
Learning joint latent representations for images and language
… Learning the joint latent representations for images and language is vital to solving many image-text tasks, including image-sentence retrieval, visual grounding, and image captioning, etc. In this thesis, we first propose two-branch neural networks for learning the similarity between these two …
-
VLEO-Bench: A Framework to Evaluate Vision-Language Models for Earth Observation Applications
… unclear to what extent capabilities on natural images transfer to Earth observation (EO) data, which are predominantly satellite and aerial images less common in VLM training data. In this work, we propose VLEO-Bench, a comprehensive evaluation framework to quantify the progress of VLMs toward …
-
Topic-Based Video Classification and Retrieval Using Machine Learning
… primarily concentrate on object detection, image classification, and image captioning. However, very little work has been shown in DL-based video-content analysis and retrieval. Due to the complex nature of time relevant information in a sequence of video frames, understanding video contents …
-
Evaluating visually grounded language capabilities using microworlds
… a new approach to assess generative tasks like image captioning.
-
Evaluating Natural Language Generation Tasks for Grammaticality, Faithfulness and Diversity
… generation tasks are explored: synthetic image captioning, football highlight generation from match statistics, and topic-shift dialogue generation. These tasks are deliberately chosen to cover a diverse range of generation scenarios. Each task provides unique grounding information and …
-
Going Deeper with Images and Natural Language
… in which machines now can categorize images into multiple classes, and detect various objects within an image, with an ability that is competitive with or even surpasses that of humans. Meanwhile, we also have witnessed similar strides in natural language processing (NLP). It is quite …
-
Evaluating interactions and decision-making by annotators and users in computer vision systems
… helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the …
-
Learning to map between domains
… something seen before. Similarly, in the medical image field and radiological science, tens of thousands of medical images (MRI, CT, etc) of patients are taken. These medical images need to be studied and interpreted. In this dissertation, we investigate a number of data-driven approaches for …
-
Natural language image description: data, models, and evaluation
Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: …