{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/100977"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/100977","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Grounding natural language phrases in images and video","abstract":"Grounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions. I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query. To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task. These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network. I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image. Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.","abstract_html":"Grounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions. I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query. To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task. These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network. I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image. Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.","abstract_has_math":false,"creators":["Plummer, Bryan A."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana","Hockenmaier, Julia","Hoiem, Derek","Brown, Matthew"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:27:05Z","date_published":"2018-09-04T20:27:05Z","updated_at":"2026-07-22T22:24:38Z","subjects":["Computer Vision, Natural Language Processing, Phrase Grounding"],"languages":["en"],"rights":["Copyright 2018 Bryan A. Plummer"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/100977","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana","Hockenmaier, Julia","Hoiem, Derek","Brown, Matthew"]},{"key":"dc:creator","label":"Author","values":["Plummer, Bryan A."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:27:05Z","2018-04-16","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision, Natural Language Processing, Phrase Grounding"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Bryan A. Plummer"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/100977"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Grounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions. I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query. To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task. These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network. I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image. Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms","The student, Bryan Plummer, accepted the attached license on 2018-04-14 at 05:42.","The student, Bryan Plummer, submitted this Dissertation for approval on 2018-04-14 at 05:57.","This Dissertation was approved for publication on 2018-04-16 at 11:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12248 on 2018-08-31 at 17:12:23","Made available in DSpace on 2018-09-04T20:27:05Z (GMT). No. of bitstreams: 3 PLUMMER-DISSERTATION-2018.pdf: 9214167 bytes, checksum: 7c4ec584f104360a0a99ea9a290b5a09 (MD5) grounding-natural-language.zip: 8908155 bytes, checksum: 79b1f7c90322e047fc2559c7486ab3c1 (MD5) LICENSE.txt: 4210 bytes, checksum: 398aac9d4dcb281c202797929dc1603f (MD5) Previous issue date: 2018-04-16"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Grounding natural language phrases in images and video"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana","Hockenmaier, Julia","Hoiem, Derek","Brown, Matthew"],"dc:creator":["Plummer, Bryan A."],"dc:date":["2018-09-04T20:27:05Z","2018-04-16","2018-05"],"dc:description":["Grounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions. I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query. To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task. These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network. I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image. Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms","The student, Bryan Plummer, accepted the attached license on 2018-04-14 at 05:42.","The student, Bryan Plummer, submitted this Dissertation for approval on 2018-04-14 at 05:57.","This Dissertation was approved for publication on 2018-04-16 at 11:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12248 on 2018-08-31 at 17:12:23","Made available in DSpace on 2018-09-04T20:27:05Z (GMT). No. of bitstreams: 3 PLUMMER-DISSERTATION-2018.pdf: 9214167 bytes, checksum: 7c4ec584f104360a0a99ea9a290b5a09 (MD5) grounding-natural-language.zip: 8908155 bytes, checksum: 79b1f7c90322e047fc2559c7486ab3c1 (MD5) LICENSE.txt: 4210 bytes, checksum: 398aac9d4dcb281c202797929dc1603f (MD5) Previous issue date: 2018-04-16"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/100977"],"dc:language":["en"],"dc:rights":["Copyright 2018 Bryan A. Plummer"],"dc:subject":["Computer Vision, Natural Language Processing, Phrase Grounding"],"dc:title":["Grounding natural language phrases in images and video"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}