{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101314"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101314","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning and adapting visual models for multiple specialized tasks","abstract":"A key requirement for any agent that wishes to interact with the visual world is the ability to understand the behavior of objects in the scene, primarily through visual means. We humans, through our cognitive system, are able to localize other people and objects in scenes, understand their relationship to the surrounding environment, and reason about not only their actions and attributes, but also about concepts which require knowledge beyond what is afforded by the pixels in visual input, such as possible future states, motion, a person’s motivations, and so on. In this thesis, we outline work that takes small steps towards solving this daunting task of replicating the human visual cognitive system. This dissertation presents methods for predicting actions, interactions with objects, and increasingly structured scenarios from single images. We devise simple methods that make use of a variety of cues by taking into account the structure inherent in the tasks we aim to solve. We show that by solving these tasks as an intermediate step and using their outputs as features, we can develop methods that operate on visual and language inputs to improve performance on tasks that require high-level image information, such as answering questions about images and producing captions for images. One issue that accompanies the learning of multiple tasks with separate deep networks, such as the work described above, is the need to store separate models, which increases storage requirements and affects scalability. We formulate and present two novel methods that draw inspiration from network pruning and weight quantization that can reuse parts of an existing network for learning new tasks with minimal additional overhead, without hurting performance on tasks that were learned earlier.","abstract_html":"A key requirement for any agent that wishes to interact with the visual world is the ability to understand the behavior of objects in the scene, primarily through visual means. We humans, through our cognitive system, are able to localize other people and objects in scenes, understand their relationship to the surrounding environment, and reason about not only their actions and attributes, but also about concepts which require knowledge beyond what is afforded by the pixels in visual input, such as possible future states, motion, a person’s motivations, and so on. In this thesis, we outline work that takes small steps towards solving this daunting task of replicating the human visual cognitive system. This dissertation presents methods for predicting actions, interactions with objects, and increasingly structured scenarios from single images. We devise simple methods that make use of a variety of cues by taking into account the structure inherent in the tasks we aim to solve. We show that by solving these tasks as an intermediate step and using their outputs as features, we can develop methods that operate on visual and language inputs to improve performance on tasks that require high-level image information, such as answering questions about images and producing captions for images. One issue that accompanies the learning of multiple tasks with separate deep networks, such as the work described above, is the need to store separate models, which increases storage requirements and affects scalability. We formulate and present two novel methods that draw inspiration from network pruning and weight quantization that can reuse parts of an existing network for learning new tasks with minimal additional overhead, without hurting performance on tasks that were learned earlier.","abstract_has_math":false,"creators":["Mallya, Arun Mohanray"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Shakhnarovich, Gregory"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:47:14Z","date_published":"2018-09-04T20:47:14Z","updated_at":"2026-07-22T22:24:38Z","subjects":["Action Recognition, Visual Relationship Detection, Image Situations, Multi Task Training"],"languages":["en"],"rights":["Copyright 2018 Arun Mallya"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101314","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Shakhnarovich, Gregory"]},{"key":"dc:creator","label":"Author","values":["Mallya, Arun Mohanray"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:47:14Z","2020-09-05T09:15:16Z","2018-04-15","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Action Recognition, Visual Relationship Detection, Image Situations, Multi Task Training"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Arun Mallya"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101314"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A key requirement for any agent that wishes to interact with the visual world is the ability to understand the behavior of objects in the scene, primarily through visual means. We humans, through our cognitive system, are able to localize other people and objects in scenes, understand their relationship to the surrounding environment, and reason about not only their actions and attributes, but also about concepts which require knowledge beyond what is afforded by the pixels in visual input, such as possible future states, motion, a person’s motivations, and so on. In this thesis, we outline work that takes small steps towards solving this daunting task of replicating the human visual cognitive system. This dissertation presents methods for predicting actions, interactions with objects, and increasingly structured scenarios from single images. We devise simple methods that make use of a variety of cues by taking into account the structure inherent in the tasks we aim to solve. We show that by solving these tasks as an intermediate step and using their outputs as features, we can develop methods that operate on visual and language inputs to improve performance on tasks that require high-level image information, such as answering questions about images and producing captions for images. One issue that accompanies the learning of multiple tasks with separate deep networks, such as the work described above, is the need to store separate models, which increases storage requirements and affects scalability. We formulate and present two novel methods that draw inspiration from network pruning and weight quantization that can reuse parts of an existing network for learning new tasks with minimal additional overhead, without hurting performance on tasks that were learned earlier.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-05-01","The student, Arun Mallya, accepted the attached license on 2018-04-13 at 16:21.","The student, Arun Mallya, submitted this Dissertation for approval on 2018-04-13 at 16:31.","This Dissertation was approved for publication on 2018-04-15 at 10:05.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12243 on 2018-08-31 at 17:28:53","Made available in DSpace on 2018-09-04T20:47:14Z (GMT). No. of bitstreams: 2 MALLYA-DISSERTATION-2018.pdf: 49245720 bytes, checksum: f6ea9de505f288ff6b60589064758518 (MD5) LICENSE.txt: 4208 bytes, checksum: 1b480cee3dd35e59d8f44090637e399b (MD5) Previous issue date: 2018-04-15","Embargo set by: Seth Robbins for item 107399 Lift date: 2020-09-04T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107399 Lift date: 2020-09-04T20:50:11Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 107399 on 2020-09-05T09:15:16Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning and adapting visual models for multiple specialized tasks"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Shakhnarovich, Gregory"],"dc:creator":["Mallya, Arun Mohanray"],"dc:date":["2018-09-04T20:47:14Z","2020-09-05T09:15:16Z","2018-04-15","2018-05"],"dc:description":["A key requirement for any agent that wishes to interact with the visual world is the ability to understand the behavior of objects in the scene, primarily through visual means. We humans, through our cognitive system, are able to localize other people and objects in scenes, understand their relationship to the surrounding environment, and reason about not only their actions and attributes, but also about concepts which require knowledge beyond what is afforded by the pixels in visual input, such as possible future states, motion, a person’s motivations, and so on. In this thesis, we outline work that takes small steps towards solving this daunting task of replicating the human visual cognitive system. This dissertation presents methods for predicting actions, interactions with objects, and increasingly structured scenarios from single images. We devise simple methods that make use of a variety of cues by taking into account the structure inherent in the tasks we aim to solve. We show that by solving these tasks as an intermediate step and using their outputs as features, we can develop methods that operate on visual and language inputs to improve performance on tasks that require high-level image information, such as answering questions about images and producing captions for images. One issue that accompanies the learning of multiple tasks with separate deep networks, such as the work described above, is the need to store separate models, which increases storage requirements and affects scalability. We formulate and present two novel methods that draw inspiration from network pruning and weight quantization that can reuse parts of an existing network for learning new tasks with minimal additional overhead, without hurting performance on tasks that were learned earlier.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-05-01","The student, Arun Mallya, accepted the attached license on 2018-04-13 at 16:21.","The student, Arun Mallya, submitted this Dissertation for approval on 2018-04-13 at 16:31.","This Dissertation was approved for publication on 2018-04-15 at 10:05.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12243 on 2018-08-31 at 17:28:53","Made available in DSpace on 2018-09-04T20:47:14Z (GMT). No. of bitstreams: 2 MALLYA-DISSERTATION-2018.pdf: 49245720 bytes, checksum: f6ea9de505f288ff6b60589064758518 (MD5) LICENSE.txt: 4208 bytes, checksum: 1b480cee3dd35e59d8f44090637e399b (MD5) Previous issue date: 2018-04-15","Embargo set by: Seth Robbins for item 107399 Lift date: 2020-09-04T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107399 Lift date: 2020-09-04T20:50:11Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 107399 on 2020-09-05T09:15:16Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101314"],"dc:language":["en"],"dc:rights":["Copyright 2018 Arun Mallya"],"dc:subject":["Action Recognition, Visual Relationship Detection, Image Situations, Multi Task Training"],"dc:title":["Learning and adapting visual models for multiple specialized tasks"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}