{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101072"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101072","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Entity-based scene understanding","abstract":"Unifying multiple descriptions to determine the details of an everyday event can be a challenging task for humans. Though incorporating other modalities like images or videos can help humans unify such descriptions, this remains a challenging task for computational systems. We define entity-based scene understanding as the task of identifying the entities in a visual scene from multiple descriptions. This task subsumes coreference resolution, bridging resolution, and grounding to produce mutually consistent relations between entity mentions and groundings between mentions and image regions. Using neural classifiers and integer linear program inference, we show that grounding is improved when forced to conform to relation predictions. We introduce the Flickr30k Entities v2 dataset, and show how our methods can be used to automatically generate similarly rich annotations for the MSCOCO dataset.","abstract_html":"Unifying multiple descriptions to determine the details of an everyday event can be a challenging task for humans. Though incorporating other modalities like images or videos can help humans unify such descriptions, this remains a challenging task for computational systems. We define entity-based scene understanding as the task of identifying the entities in a visual scene from multiple descriptions. This task subsumes coreference resolution, bridging resolution, and grounding to produce mutually consistent relations between entity mentions and groundings between mentions and image regions. Using neural classifiers and integer linear program inference, we show that grounding is improved when forced to conform to relation predictions. We introduce the Flickr30k Entities v2 dataset, and show how our methods can be used to automatically generate similarly rich annotations for the MSCOCO dataset.","abstract_has_math":false,"creators":["Cervantes, Christopher Michael"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hockenmaier, Julia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:31:56Z","date_published":"2018-09-04T20:31:56Z","updated_at":"2026-07-22T22:24:38Z","subjects":["coreference, bridging, grounding, neural networks, LSTM, Flickr30k Entities, MSCOCO"],"languages":["en"],"rights":["Copyright 2018 Christopher Cervantes"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101072","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hockenmaier, Julia"]},{"key":"dc:creator","label":"Author","values":["Cervantes, Christopher Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:31:56Z","2018-04-25","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["coreference, bridging, grounding, neural networks, LSTM, Flickr30k Entities, MSCOCO"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Christopher Cervantes"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101072"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Unifying multiple descriptions to determine the details of an everyday event can be a challenging task for humans. Though incorporating other modalities like images or videos can help humans unify such descriptions, this remains a challenging task for computational systems. We define entity-based scene understanding as the task of identifying the entities in a visual scene from multiple descriptions. This task subsumes coreference resolution, bridging resolution, and grounding to produce mutually consistent relations between entity mentions and groundings between mentions and image regions. Using neural classifiers and integer linear program inference, we show that grounding is improved when forced to conform to relation predictions. We introduce the Flickr30k Entities v2 dataset, and show how our methods can be used to automatically generate similarly rich annotations for the MSCOCO dataset.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms","The student, Christopher Cervantes, accepted the attached license on 2018-04-25 at 11:34.","The student, Christopher Cervantes, submitted this Thesis for approval on 2018-04-25 at 11:49.","This Thesis was approved for publication on 2018-04-25 at 16:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12478 on 2018-08-31 at 17:14:40","Made available in DSpace on 2018-09-04T20:31:56Z (GMT). No. of bitstreams: 3 CERVANTES-THESIS-2018.pdf: 5428393 bytes, checksum: a72b3a67170ff3299c870464f1ae331b (MD5) source.tar.gz: 5899853 bytes, checksum: e42acbb834091ab950311023ccaa5a5e (MD5) LICENSE.txt: 4218 bytes, checksum: 24b8fee58464e584614d5f0e2fe4371d (MD5) Previous issue date: 2018-04-25"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Entity-based scene understanding"]}]}],"canonical_facts":{"dc:contributor":["Hockenmaier, Julia"],"dc:creator":["Cervantes, Christopher Michael"],"dc:date":["2018-09-04T20:31:56Z","2018-04-25","2018-05"],"dc:description":["Unifying multiple descriptions to determine the details of an everyday event can be a challenging task for humans. Though incorporating other modalities like images or videos can help humans unify such descriptions, this remains a challenging task for computational systems. We define entity-based scene understanding as the task of identifying the entities in a visual scene from multiple descriptions. This task subsumes coreference resolution, bridging resolution, and grounding to produce mutually consistent relations between entity mentions and groundings between mentions and image regions. Using neural classifiers and integer linear program inference, we show that grounding is improved when forced to conform to relation predictions. We introduce the Flickr30k Entities v2 dataset, and show how our methods can be used to automatically generate similarly rich annotations for the MSCOCO dataset.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms","The student, Christopher Cervantes, accepted the attached license on 2018-04-25 at 11:34.","The student, Christopher Cervantes, submitted this Thesis for approval on 2018-04-25 at 11:49.","This Thesis was approved for publication on 2018-04-25 at 16:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12478 on 2018-08-31 at 17:14:40","Made available in DSpace on 2018-09-04T20:31:56Z (GMT). No. of bitstreams: 3 CERVANTES-THESIS-2018.pdf: 5428393 bytes, checksum: a72b3a67170ff3299c870464f1ae331b (MD5) source.tar.gz: 5899853 bytes, checksum: e42acbb834091ab950311023ccaa5a5e (MD5) LICENSE.txt: 4218 bytes, checksum: 24b8fee58464e584614d5f0e2fe4371d (MD5) Previous issue date: 2018-04-25"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101072"],"dc:language":["en"],"dc:rights":["Copyright 2018 Christopher Cervantes"],"dc:subject":["coreference, bridging, grounding, neural networks, LSTM, Flickr30k Entities, MSCOCO"],"dc:title":["Entity-based scene understanding"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}