{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/102864"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/102864","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Weakly supervised learning from referring expression: Challenge and directions","abstract":"We explore methods of weakly supervised learning from referring expression. Unlike traditional fully supervised semantic segmentation of object recognition tasks, in which a a small set of discrete class bases is provided, the referring expression task is performed associated with a sentence phrase, e.g. “the dude on the dolphin”. Previous approaches use LSTM and fully convolutional network and have fairly good results under fully supervised setting. However, the fully supervised setting is limited by manual labeling of segmentation masks, which requires a significant amount of human labor. Therefore, we work on an approach to perform segmentation with only image level language descriptions. Under our weakly supervised setting, we are only provided with input images and the corresponding sentence descriptions, without the pixel level labeling for each image as ground truth. In order to get supervision only from language description, we utilize the multiple instance learning loss. We first develop an end-to-end model to localize the image content corresponding to the language expressions. In this model, we use GloVe and ELMo sentence embeddings to get a vector representation for each sentence and combined with image features from a fully convolutional network. However, the sentence level model is hard to interpret hence we also study a more fundamental problem of weakly supervised object localization from referring expressions. We compare the performance of the sentence level model on this task to an alternative word-level model. Our investigation suggests that breaking the referring expressions localization problem into smaller more manageable components is promising.","abstract_html":"We explore methods of weakly supervised learning from referring expression. Unlike traditional fully supervised semantic segmentation of object recognition tasks, in which a a small set of discrete class bases is provided, the referring expression task is performed associated with a sentence phrase, e.g. “the dude on the dolphin”. Previous approaches use LSTM and fully convolutional network and have fairly good results under fully supervised setting. However, the fully supervised setting is limited by manual labeling of segmentation masks, which requires a significant amount of human labor. Therefore, we work on an approach to perform segmentation with only image level language descriptions. Under our weakly supervised setting, we are only provided with input images and the corresponding sentence descriptions, without the pixel level labeling for each image as ground truth. In order to get supervision only from language description, we utilize the multiple instance learning loss. We first develop an end-to-end model to localize the image content corresponding to the language expressions. In this model, we use GloVe and ELMo sentence embeddings to get a vector representation for each sentence and combined with image features from a fully convolutional network. However, the sentence level model is hard to interpret hence we also study a more fundamental problem of weakly supervised object localization from referring expressions. We compare the performance of the sentence level model on this task to an alternative word-level model. Our investigation suggests that breaking the referring expressions localization problem into smaller more manageable components is promising.","abstract_has_math":false,"creators":["Dong, Taiyu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hoiem, Derek"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-02-07T20:44:30Z","date_published":"2019-02-07T20:44:30Z","updated_at":"2026-07-22T22:24:42Z","subjects":["weakly supervised learning","object localization","segmentation from natural language","object localization from nature language"],"languages":["en"],"rights":["Copyright 2018 Taiyu Dong"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/102864","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hoiem, Derek"]},{"key":"dc:creator","label":"Author","values":["Dong, Taiyu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-02-07T20:44:30Z","2021-02-08T10:15:18Z","2018-12-13","2018-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["weakly supervised learning","object localization","segmentation from natural language","object localization from nature language"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Taiyu Dong"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/102864"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We explore methods of weakly supervised learning from referring expression. Unlike traditional fully supervised semantic segmentation of object recognition tasks, in which a a small set of discrete class bases is provided, the referring expression task is performed associated with a sentence phrase, e.g. “the dude on the dolphin”. Previous approaches use LSTM and fully convolutional network and have fairly good results under fully supervised setting. However, the fully supervised setting is limited by manual labeling of segmentation masks, which requires a significant amount of human labor. Therefore, we work on an approach to perform segmentation with only image level language descriptions. Under our weakly supervised setting, we are only provided with input images and the corresponding sentence descriptions, without the pixel level labeling for each image as ground truth. In order to get supervision only from language description, we utilize the multiple instance learning loss. We first develop an end-to-end model to localize the image content corresponding to the language expressions. In this model, we use GloVe and ELMo sentence embeddings to get a vector representation for each sentence and combined with image features from a fully convolutional network. However, the sentence level model is hard to interpret hence we also study a more fundamental problem of weakly supervised object localization from referring expressions. We compare the performance of the sentence level model on this task to an alternative word-level model. Our investigation suggests that breaking the referring expressions localization problem into smaller more manageable components is promising.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-12-01","The student, Taiyu Dong, accepted the attached license on 2018-12-13 at 10:16.","The student, Taiyu Dong, submitted this Thesis for approval on 2018-12-13 at 10:25.","This Thesis was approved for publication on 2018-12-13 at 14:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13313 on 2019-02-07 at 14:23:34","Made available in DSpace on 2019-02-07T20:44:30Z (GMT). No. of bitstreams: 2 DONG-THESIS-2018.pdf: 4913301 bytes, checksum: 9cdc521def4fb7ffb6a0bec7153a59d9 (MD5) LICENSE.txt: 4207 bytes, checksum: b9b9ba6876dee8693c6e213ae5d64f5b (MD5) Previous issue date: 2018-12-13","Embargo set by: Seth Robbins for item 109890 Lift date: 2021-02-07T20:44:35Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 109890 on 2021-02-08T10:15:18Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Weakly supervised learning from referring expression: Challenge and directions"]}]}],"canonical_facts":{"dc:contributor":["Hoiem, Derek"],"dc:creator":["Dong, Taiyu"],"dc:date":["2019-02-07T20:44:30Z","2021-02-08T10:15:18Z","2018-12-13","2018-12"],"dc:description":["We explore methods of weakly supervised learning from referring expression. Unlike traditional fully supervised semantic segmentation of object recognition tasks, in which a a small set of discrete class bases is provided, the referring expression task is performed associated with a sentence phrase, e.g. “the dude on the dolphin”. Previous approaches use LSTM and fully convolutional network and have fairly good results under fully supervised setting. However, the fully supervised setting is limited by manual labeling of segmentation masks, which requires a significant amount of human labor. Therefore, we work on an approach to perform segmentation with only image level language descriptions. Under our weakly supervised setting, we are only provided with input images and the corresponding sentence descriptions, without the pixel level labeling for each image as ground truth. In order to get supervision only from language description, we utilize the multiple instance learning loss. We first develop an end-to-end model to localize the image content corresponding to the language expressions. In this model, we use GloVe and ELMo sentence embeddings to get a vector representation for each sentence and combined with image features from a fully convolutional network. However, the sentence level model is hard to interpret hence we also study a more fundamental problem of weakly supervised object localization from referring expressions. We compare the performance of the sentence level model on this task to an alternative word-level model. Our investigation suggests that breaking the referring expressions localization problem into smaller more manageable components is promising.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-12-01","The student, Taiyu Dong, accepted the attached license on 2018-12-13 at 10:16.","The student, Taiyu Dong, submitted this Thesis for approval on 2018-12-13 at 10:25.","This Thesis was approved for publication on 2018-12-13 at 14:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13313 on 2019-02-07 at 14:23:34","Made available in DSpace on 2019-02-07T20:44:30Z (GMT). No. of bitstreams: 2 DONG-THESIS-2018.pdf: 4913301 bytes, checksum: 9cdc521def4fb7ffb6a0bec7153a59d9 (MD5) LICENSE.txt: 4207 bytes, checksum: b9b9ba6876dee8693c6e213ae5d64f5b (MD5) Previous issue date: 2018-12-13","Embargo set by: Seth Robbins for item 109890 Lift date: 2021-02-07T20:44:35Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 109890 on 2021-02-08T10:15:18Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/102864"],"dc:language":["en"],"dc:rights":["Copyright 2018 Taiyu Dong"],"dc:subject":["weakly supervised learning","object localization","segmentation from natural language","object localization from nature language"],"dc:title":["Weakly supervised learning from referring expression: Challenge and directions"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}