{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/89008"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/89008","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Natural language image description: data, models, and evaluation","abstract":"Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: 8cf5a4d43e5d4e329daf3dd841c8fc71 (MD5) workshopPermission.pdf: 121420 bytes, checksum: 73f2cbfefabf39534bc95c16e1f4a7a7 (MD5) Previous issue date: 2015-11-25","abstract_html":"Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: 8cf5a4d43e5d4e329daf3dd841c8fc71 (MD5) workshopPermission.pdf: 121420 bytes, checksum: 73f2cbfefabf39534bc95c16e1f4a7a7 (MD5) Previous issue date: 2015-11-25","abstract_has_math":false,"creators":["Hodosh, Micah A"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hockenmaier, Julia","Dolan, Bill","Forsyth, David","Roth, Dan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-03-02T19:33:52Z","date_published":"2016-03-02T19:33:52Z","updated_at":"2026-07-22T22:26:32Z","subjects":["Computer Vision","Natural Language Processing","Image Description","Neural Networks","computer vision (CV)","natural language processing (NLP)","Machine Learning","Machine Learning Applications","Image Captioning"],"languages":["en"],"rights":["Copyright 2015 Micah A Hodosh"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/89008","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hockenmaier, Julia","Dolan, Bill","Forsyth, David","Roth, Dan"]},{"key":"dc:creator","label":"Author","values":["Hodosh, Micah A"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-03-02T19:33:52Z","2015-11-25","2015-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Natural Language Processing","Image Description","Neural Networks","computer vision (CV)","natural language processing (NLP)","Machine Learning","Machine Learning Applications","Image Captioning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Micah A Hodosh"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/89008"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: 8cf5a4d43e5d4e329daf3dd841c8fc71 (MD5) workshopPermission.pdf: 121420 bytes, checksum: 73f2cbfefabf39534bc95c16e1f4a7a7 (MD5) Previous issue date: 2015-11-25","Automatically describing an image with a concise natural language description is an ambitious and emerging task bringing together the Natural Language and Computer Vision communities. With any emerging task, the necessary groundwork developing appropriate datasets, strong baseline models, and evaluation frameworks is key. In this thesis, we introduce the rst large datasets speci cally designed with image description in mind, focusing on concrete descriptions that can be gleaned from the image alone. Furthermore, we develop strong baseline models that show the need to model language beyond a simple bag-of-words approach to increase performance. Most importantly, we introduce a ranking based framework for comparing image description models. We show that this framework is more reliable and accurate than the conventional wisdom of evaluating on novel model generated text. As this task has gained popularity recently, we further analyze the drawbacks of current evaluation methods, and put forth concrete extensions to our ranking framework that will guide progress towards modeling the association of natural language and the images the language describes.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Micah Hodosh, accepted the attached license on 2015-11-24 at 18:30.","The student, Micah Hodosh, submitted this Dissertation for approval on 2015-11-24 at 20:02.","This Dissertation was approved for publication on 2015-11-25 at 10:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8838 on 2016-03-02 at 12:50:24"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Natural language image description: data, models, and evaluation"]}]}],"canonical_facts":{"dc:contributor":["Hockenmaier, Julia","Dolan, Bill","Forsyth, David","Roth, Dan"],"dc:creator":["Hodosh, Micah A"],"dc:date":["2016-03-02T19:33:52Z","2015-11-25","2015-12"],"dc:description":["Made available in DSpace on 2016-03-02T19:33:52Z (GMT). No. of bitstreams: 4 HODOSH-DISSERTATION-2015.pdf: 33671796 bytes, checksum: 1f9350bc33a01a78722da502ecf52e0a (MD5) JAIR Permission.pdf: 184972 bytes, checksum: c985c2684d03383ef5292dbf559e1465 (MD5) LICENSE.txt: 4209 bytes, checksum: 8cf5a4d43e5d4e329daf3dd841c8fc71 (MD5) workshopPermission.pdf: 121420 bytes, checksum: 73f2cbfefabf39534bc95c16e1f4a7a7 (MD5) Previous issue date: 2015-11-25","Automatically describing an image with a concise natural language description is an ambitious and emerging task bringing together the Natural Language and Computer Vision communities. With any emerging task, the necessary groundwork developing appropriate datasets, strong baseline models, and evaluation frameworks is key. In this thesis, we introduce the rst large datasets speci cally designed with image description in mind, focusing on concrete descriptions that can be gleaned from the image alone. Furthermore, we develop strong baseline models that show the need to model language beyond a simple bag-of-words approach to increase performance. Most importantly, we introduce a ranking based framework for comparing image description models. We show that this framework is more reliable and accurate than the conventional wisdom of evaluating on novel model generated text. As this task has gained popularity recently, we further analyze the drawbacks of current evaluation methods, and put forth concrete extensions to our ranking framework that will guide progress towards modeling the association of natural language and the images the language describes.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Micah Hodosh, accepted the attached license on 2015-11-24 at 18:30.","The student, Micah Hodosh, submitted this Dissertation for approval on 2015-11-24 at 20:02.","This Dissertation was approved for publication on 2015-11-25 at 10:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8838 on 2016-03-02 at 12:50:24"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/89008"],"dc:language":["en"],"dc:rights":["Copyright 2015 Micah A Hodosh"],"dc:subject":["Computer Vision","Natural Language Processing","Image Description","Neural Networks","computer vision (CV)","natural language processing (NLP)","Machine Learning","Machine Learning Applications","Image Captioning"],"dc:title":["Natural language image description: data, models, and evaluation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:32Z"}