{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/46633"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/46633","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Using the visual denotations of image captions for semantic inference","abstract":"Semantic inference is essential to natural language understanding. There are two different traditional approaches to semantic inference. The logic-based approach translates utterances into a formal meaning representation that is amenable to logical proofs. The vector-based approach maps words to vectors that are based on the contexts in which the words appear in utterances. Real-valued similarities are used in place of logical inferences. We introduce the notion of the visual denotation of an utterance, which is the set of images that it describes. This notion borrows the abstract concept of a denotation of an utterance as the set of possible worlds in which the utterance is true from the logic-based approach, and instantiates possible worlds as images. In this dissertation, we also show how visual denotations can be created for descriptions of everyday entities and events. Additionally, we demonstrate that visual denotations can be used as a new model of semantic similarity, and that this model is better at identifying entailment relations between descriptions of images than traditional distributional similarities. In order to do this, we create an image caption corpus consisting of captions and images depicting everyday actions. This corpus has a number of useful features that would assist in investigating everyday events and the different ways in which they can be described. We use the captions in the corpus as the starting point for producing caption fragments with larger visual denotations. We accomplish that by creating a denotation graph, a subsumption hierarchy over the captions that links captions and images that depict them, that also allows for the visualization and navigation of the image caption corpus in an intuitive manner.","abstract_html":"Semantic inference is essential to natural language understanding. There are two different traditional approaches to semantic inference. The logic-based approach translates utterances into a formal meaning representation that is amenable to logical proofs. The vector-based approach maps words to vectors that are based on the contexts in which the words appear in utterances. Real-valued similarities are used in place of logical inferences. We introduce the notion of the visual denotation of an utterance, which is the set of images that it describes. This notion borrows the abstract concept of a denotation of an utterance as the set of possible worlds in which the utterance is true from the logic-based approach, and instantiates possible worlds as images. In this dissertation, we also show how visual denotations can be created for descriptions of everyday entities and events. Additionally, we demonstrate that visual denotations can be used as a new model of semantic similarity, and that this model is better at identifying entailment relations between descriptions of images than traditional distributional similarities. In order to do this, we create an image caption corpus consisting of captions and images depicting everyday actions. This corpus has a number of useful features that would assist in investigating everyday events and the different ways in which they can be described. We use the captions in the corpus as the starting point for producing caption fragments with larger visual denotations. We accomplish that by creating a denotation graph, a subsumption hierarchy over the captions that links captions and images that depict them, that also allows for the visualization and navigation of the image caption corpus in an intuitive manner.","abstract_has_math":false,"creators":["Young, Peter"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hockenmaier, Julia C.","DeJong, Gerald F.","Palmer, Martha","Roth, Dan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-01-16T17:56:53Z","date_published":"2014-01-16T17:56:53Z","updated_at":"2026-07-22T22:25:36Z","subjects":["visual denotation","natural language processing (nlp)","image caption corpus","denotation graph"],"languages":["en"],"rights":["Copyright 2013 Peter Young"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/46633","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hockenmaier, Julia C.","DeJong, Gerald F.","Palmer, Martha","Roth, Dan"]},{"key":"dc:creator","label":"Author","values":["Young, Peter"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-01-16T17:56:53Z","2013-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["visual denotation","natural language processing (nlp)","image caption corpus","denotation graph"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Peter Young"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/46633"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Semantic inference is essential to natural language understanding. There are two different traditional approaches to semantic inference. The logic-based approach translates utterances into a formal meaning representation that is amenable to logical proofs. The vector-based approach maps words to vectors that are based on the contexts in which the words appear in utterances. Real-valued similarities are used in place of logical inferences. We introduce the notion of the visual denotation of an utterance, which is the set of images that it describes. This notion borrows the abstract concept of a denotation of an utterance as the set of possible worlds in which the utterance is true from the logic-based approach, and instantiates possible worlds as images. In this dissertation, we also show how visual denotations can be created for descriptions of everyday entities and events. Additionally, we demonstrate that visual denotations can be used as a new model of semantic similarity, and that this model is better at identifying entailment relations between descriptions of images than traditional distributional similarities. In order to do this, we create an image caption corpus consisting of captions and images depicting everyday actions. This corpus has a number of useful features that would assist in investigating everyday events and the different ways in which they can be described. We use the captions in the corpus as the starting point for producing caption fragments with larger visual denotations. We accomplish that by creating a denotation graph, a subsumption hierarchy over the captions that links captions and images that depict them, that also allows for the visualization and navigation of the image caption corpus in an intuitive manner.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-12-04T19:54:56Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Young_Peter.pdf: 2992405 bytes, checksum: bc161d17492e6c6a4e8c44eddc12b90a (MD5) Young_Peter.pdf: 2992396 bytes, checksum: 27bc8fbd147be0a3c7859755b85f6e98 (MD5)","Made available in DSpace on 2014-01-16T17:56:53Z (GMT). No. of bitstreams: 2 Peter_Young.pdf: 2992396 bytes, checksum: 27bc8fbd147be0a3c7859755b85f6e98 (MD5) license.txt: 4060 bytes, checksum: 36fb7660ab19386b6956de38ba6f4f97 (MD5)"]},{"key":"dc:title","label":"Title","values":["Using the visual denotations of image captions for semantic inference"]}]}],"canonical_facts":{"dc:contributor":["Hockenmaier, Julia C.","DeJong, Gerald F.","Palmer, Martha","Roth, Dan"],"dc:creator":["Young, Peter"],"dc:date":["2014-01-16T17:56:53Z","2013-12"],"dc:description":["Semantic inference is essential to natural language understanding. There are two different traditional approaches to semantic inference. The logic-based approach translates utterances into a formal meaning representation that is amenable to logical proofs. The vector-based approach maps words to vectors that are based on the contexts in which the words appear in utterances. Real-valued similarities are used in place of logical inferences. We introduce the notion of the visual denotation of an utterance, which is the set of images that it describes. This notion borrows the abstract concept of a denotation of an utterance as the set of possible worlds in which the utterance is true from the logic-based approach, and instantiates possible worlds as images. In this dissertation, we also show how visual denotations can be created for descriptions of everyday entities and events. Additionally, we demonstrate that visual denotations can be used as a new model of semantic similarity, and that this model is better at identifying entailment relations between descriptions of images than traditional distributional similarities. In order to do this, we create an image caption corpus consisting of captions and images depicting everyday actions. This corpus has a number of useful features that would assist in investigating everyday events and the different ways in which they can be described. We use the captions in the corpus as the starting point for producing caption fragments with larger visual denotations. We accomplish that by creating a denotation graph, a subsumption hierarchy over the captions that links captions and images that depict them, that also allows for the visualization and navigation of the image caption corpus in an intuitive manner.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-12-04T19:54:56Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Young_Peter.pdf: 2992405 bytes, checksum: bc161d17492e6c6a4e8c44eddc12b90a (MD5) Young_Peter.pdf: 2992396 bytes, checksum: 27bc8fbd147be0a3c7859755b85f6e98 (MD5)","Made available in DSpace on 2014-01-16T17:56:53Z (GMT). No. of bitstreams: 2 Peter_Young.pdf: 2992396 bytes, checksum: 27bc8fbd147be0a3c7859755b85f6e98 (MD5) license.txt: 4060 bytes, checksum: 36fb7660ab19386b6956de38ba6f4f97 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/46633"],"dc:language":["en"],"dc:rights":["Copyright 2013 Peter Young"],"dc:subject":["visual denotation","natural language processing (nlp)","image caption corpus","denotation graph"],"dc:title":["Using the visual denotations of image captions for semantic inference"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:36Z"}