University of Illinois at Urbana-Champaign
Grounding natural language phrases in images and video
Abstract
dc:descriptionGrounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions. I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query. To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task. These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network. I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image. Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2018
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Plummer, Bryan A.
- Contributors dc:contributor
-
- Lazebnik, Svetlana
- Hockenmaier, Julia
- Hoiem, Derek
- Brown, Matthew
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- Copyright 2018 Bryan A. Plummer
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/100977
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/100977