{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/98359"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/98359","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning visual tasks with selective attention","abstract":"Knowing where to look in an image can significantly improve performance in computer vision tasks by eliminating irrelevant information from the rest of the input image, and by breaking down complex scenes into simpler and more familiar sub-components. We show that a framework for identifying multiple task-relevant regions can be learned in current state-of-the-art deep network architectures, resulting in significant gains in several visual prediction tasks. We will demonstrate both directly and indirectly supervised models for selecting image regions and show how they can improve performance over baselines by means of focusing on the right areas.","abstract_html":"Knowing where to look in an image can significantly improve performance in computer vision tasks by eliminating irrelevant information from the rest of the input image, and by breaking down complex scenes into simpler and more familiar sub-components. We show that a framework for identifying multiple task-relevant regions can be learned in current state-of-the-art deep network architectures, resulting in significant gains in several visual prediction tasks. We will demonstrate both directly and indirectly supervised models for selecting image regions and show how they can improve performance over baselines by means of focusing on the right areas.","abstract_has_math":false,"creators":["Shih, Kevin Jonathan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hoiem, Derek","Lazebnik, Svetlana","Forsyth, David","Parikh, Devi"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-09-29T17:56:40Z","date_published":"2017-09-29T17:56:40Z","updated_at":"2026-07-22T22:24:35Z","subjects":["Computer vision","Visual attention","Visual question answering (VQA)","Keypoint localization","Part localization","Image recognition","Fine-grained image recognition","Deep learning","Multi-task learning","Machine learning"],"languages":["en"],"rights":["Copyright 2017 Kevin Jonathan Shih"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/98359","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hoiem, Derek","Lazebnik, Svetlana","Forsyth, David","Parikh, Devi"]},{"key":"dc:creator","label":"Author","values":["Shih, Kevin Jonathan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-09-29T17:56:40Z","2017-07-11","2017-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer vision","Visual attention","Visual question answering (VQA)","Keypoint localization","Part localization","Image recognition","Fine-grained image recognition","Deep learning","Multi-task learning","Machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Kevin Jonathan Shih"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/98359"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Knowing where to look in an image can significantly improve performance in computer vision tasks by eliminating irrelevant information from the rest of the input image, and by breaking down complex scenes into simpler and more familiar sub-components. We show that a framework for identifying multiple task-relevant regions can be learned in current state-of-the-art deep network architectures, resulting in significant gains in several visual prediction tasks. We will demonstrate both directly and indirectly supervised models for selecting image regions and show how they can improve performance over baselines by means of focusing on the right areas.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Kevin Shih, accepted the attached license on 2017-07-10 at 12:45.","The student, Kevin Shih, submitted this Dissertation for approval on 2017-07-10 at 13:18.","This Dissertation was approved for publication on 2017-07-11 at 15:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11368 on 2017-09-29 at 11:29:29","Made available in DSpace on 2017-09-29T17:56:40Z (GMT). No. of bitstreams: 2 SHIH-DISSERTATION-2017.pdf: 35992565 bytes, checksum: 0236e3afe4b94ec89729250662a7eb76 (MD5) LICENSE.txt: 4207 bytes, checksum: 850b64b383db31c4fc6d39801a3eab05 (MD5) Previous issue date: 2017-07-11"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning visual tasks with selective attention"]}]}],"canonical_facts":{"dc:contributor":["Hoiem, Derek","Lazebnik, Svetlana","Forsyth, David","Parikh, Devi"],"dc:creator":["Shih, Kevin Jonathan"],"dc:date":["2017-09-29T17:56:40Z","2017-07-11","2017-08"],"dc:description":["Knowing where to look in an image can significantly improve performance in computer vision tasks by eliminating irrelevant information from the rest of the input image, and by breaking down complex scenes into simpler and more familiar sub-components. We show that a framework for identifying multiple task-relevant regions can be learned in current state-of-the-art deep network architectures, resulting in significant gains in several visual prediction tasks. We will demonstrate both directly and indirectly supervised models for selecting image regions and show how they can improve performance over baselines by means of focusing on the right areas.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Kevin Shih, accepted the attached license on 2017-07-10 at 12:45.","The student, Kevin Shih, submitted this Dissertation for approval on 2017-07-10 at 13:18.","This Dissertation was approved for publication on 2017-07-11 at 15:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11368 on 2017-09-29 at 11:29:29","Made available in DSpace on 2017-09-29T17:56:40Z (GMT). No. of bitstreams: 2 SHIH-DISSERTATION-2017.pdf: 35992565 bytes, checksum: 0236e3afe4b94ec89729250662a7eb76 (MD5) LICENSE.txt: 4207 bytes, checksum: 850b64b383db31c4fc6d39801a3eab05 (MD5) Previous issue date: 2017-07-11"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/98359"],"dc:language":["en"],"dc:rights":["Copyright 2017 Kevin Jonathan Shih"],"dc:subject":["Computer vision","Visual attention","Visual question answering (VQA)","Keypoint localization","Part localization","Image recognition","Fine-grained image recognition","Deep learning","Multi-task learning","Machine learning"],"dc:title":["Learning visual tasks with selective attention"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:35Z"}