{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115410"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115410","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Collaborative embodied agents","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Jain, Unnat"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Schwing, Alexander","Lazebnik, Svetlana","Hoiem, Derek","Jiang, Nan","Grauman, Kristen"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["embodied agents","visual navigation","multi-agent systems","multi-agent RL","imitation learning","simulation","virtual environments","embodied AI"],"languages":["en","eng"],"rights":["Copyright 2022 Unnat Jain"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115410","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schwing, Alexander","Lazebnik, Svetlana","Hoiem, Derek","Jiang, Nan","Grauman, Kristen"]},{"key":"dc:creator","label":"Author","values":["Jain, Unnat"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-20"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["embodied agents","visual navigation","multi-agent systems","multi-agent RL","imitation learning","simulation","virtual environments","embodied AI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Unnat Jain"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115410"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Unnat Jain, accepted the attached license on 2022-04-20 at 14:07.","The student, Unnat Jain, submitted this Dissertation for approval on 2022-04-20 at 14:08.","This Dissertation was approved for publication on 2022-04-20 at 16:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17851 on 2022-11-11 at 13:42:34","In recent years, research in computer vision has increasingly focused on solving realistic tasks in an interactive or embodied setting. Research in embodied agents lies at the intersection of computer vision, reinforcement learning, and language understanding. Prior research and tasks (such as navigation, question answering, language grounded navigation) have equipping agents with intelligent skills with a focus on a single agent. Going forward, we believe multi-agent learning can facilitate solving increasingly complex tasks such as driving, playing sports, or moving heavy objects, or rearranging them inside the house. Taking first steps in this direction, we build AI Agents that can collaborate and communicate in virtual visual worlds. Particularly, we'll discuss collaborative learning within both homogeneous and heterogeneous sets of agents. For collaborative homogeneous agents, we include (1) formulation of collaborative tasks and effect of 'explicit' and 'implicit' communication and (2) rich 'mixture-of-marginals' policies that overcome restrictions of existing decentralized multi-agent policies. For collaborative heterogeneous agents, we include research that (3) allow learning of embodied agents from free supervision from simplistic gridworlds via a 'GridToPix' methodology and (4) how to collaborate and learn from teachers that enjoy more privilege than the student. Overall, this dissertation lays the foundations of how visual AI agents can develop skills outside the silo of learning by themselves i.e. social learning. Going forward, it would be exciting to see how to learn from other agents in your surroundings from in-the-wild videos and test sim-to-real transfer of embodied ideas and results to robots."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Collaborative embodied agents"]}]}],"canonical_facts":{"dc:contributor":["Schwing, Alexander","Lazebnik, Svetlana","Hoiem, Derek","Jiang, Nan","Grauman, Kristen"],"dc:creator":["Jain, Unnat"],"dc:date":["2022-05","2022-04-20"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Unnat Jain, accepted the attached license on 2022-04-20 at 14:07.","The student, Unnat Jain, submitted this Dissertation for approval on 2022-04-20 at 14:08.","This Dissertation was approved for publication on 2022-04-20 at 16:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17851 on 2022-11-11 at 13:42:34","In recent years, research in computer vision has increasingly focused on solving realistic tasks in an interactive or embodied setting. Research in embodied agents lies at the intersection of computer vision, reinforcement learning, and language understanding. Prior research and tasks (such as navigation, question answering, language grounded navigation) have equipping agents with intelligent skills with a focus on a single agent. Going forward, we believe multi-agent learning can facilitate solving increasingly complex tasks such as driving, playing sports, or moving heavy objects, or rearranging them inside the house. Taking first steps in this direction, we build AI Agents that can collaborate and communicate in virtual visual worlds. Particularly, we'll discuss collaborative learning within both homogeneous and heterogeneous sets of agents. For collaborative homogeneous agents, we include (1) formulation of collaborative tasks and effect of 'explicit' and 'implicit' communication and (2) rich 'mixture-of-marginals' policies that overcome restrictions of existing decentralized multi-agent policies. For collaborative heterogeneous agents, we include research that (3) allow learning of embodied agents from free supervision from simplistic gridworlds via a 'GridToPix' methodology and (4) how to collaborate and learn from teachers that enjoy more privilege than the student. Overall, this dissertation lays the foundations of how visual AI agents can develop skills outside the silo of learning by themselves i.e. social learning. Going forward, it would be exciting to see how to learn from other agents in your surroundings from in-the-wild videos and test sim-to-real transfer of embodied ideas and results to robots."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115410"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Unnat Jain"],"dc:subject":["embodied agents","visual navigation","multi-agent systems","multi-agent RL","imitation learning","simulation","virtual environments","embodied AI"],"dc:title":["Collaborative embodied agents"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}