{"id":{"repo_id":"penn","oai_identifier":"oai:repository.upenn.edu:20.500.14332/58993"},"canonical_url":"https://search.dev.ndltd.org/etd/penn/oai:repository.upenn.edu:20.500.14332/58993","repository":{"repo_id":"penn","name":"University of Pennsylvania","base_url":"https://repository.upenn.edu/server/oai/request"},"display":{"title":"LEARNING TO ACT FROM DIVERSE DATA SOURCES VIA WORLD MODELS","abstract":"The next frontier in learning to act is generalization - the ability of the agent to operate in a diverse set of environments and to solve a diverse set of tasks. How can we learn generalist agents? I argue that learning world models that predict future outcomes of actions directly from image observations is a uniquely suitable approach for training generalizable agents. I will discuss how world models can enable powerful unsupervised exploration and how to use a single world model to learn a diverse range of tasks. An implementation of this agent learns to solve a variety of kitchen or sorting tasks using a simulated robotic arm, as well as achieve arbitrary poses with a biped or a quadruped agent without any reward or demonstration supervision. I will further discuss how we can scale world model training by considering datasets of passive videos and other improvements to world models and planning. An agent that possesses a world model can use it to learn general knowledge from diverse datasets and to solve diverse tasks; I believe world models provide a principled and promising path towards building more and more general-purpose machines.","abstract_html":"The next frontier in learning to act is generalization - the ability of the agent to operate in a diverse set of environments and to solve a diverse set of tasks. How can we learn generalist agents? I argue that learning world models that predict future outcomes of actions directly from image observations is a uniquely suitable approach for training generalizable agents. I will discuss how world models can enable powerful unsupervised exploration and how to use a single world model to learn a diverse range of tasks. An implementation of this agent learns to solve a variety of kitchen or sorting tasks using a simulated robotic arm, as well as achieve arbitrary poses with a biped or a quadruped agent without any reward or demonstration supervision. I will further discuss how we can scale world model training by considering datasets of passive videos and other improvements to world models and planning. An agent that possesses a world model can use it to learn general knowledge from diverse datasets and to solve diverse tasks; I believe world models provide a principled and promising path towards building more and more general-purpose machines.","abstract_has_math":false,"creators":["Rybkin, Oleh"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Daniilidis, Kostas","Levine, Sergey"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023","date_published":"2023","updated_at":"2026-07-24T03:45:26Z","subjects":["Computer Sciences"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://repository.upenn.edu/handle/20.500.14332/58993","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Daniilidis, Kostas","Levine, Sergey"]},{"key":"dc:creator","label":"Author","values":["Rybkin, Oleh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-11-22T15:55:51Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-11-22T15:55:51Z"]},{"key":"dc:date.issued","label":"Date","values":["2023"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation/Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Sciences"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://repository.upenn.edu/handle/20.500.14332/58993"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The next frontier in learning to act is generalization - the ability of the agent to operate in a diverse set of environments and to solve a diverse set of tasks. How can we learn generalist agents? I argue that learning world models that predict future outcomes of actions directly from image observations is a uniquely suitable approach for training generalizable agents. I will discuss how world models can enable powerful unsupervised exploration and how to use a single world model to learn a diverse range of tasks. An implementation of this agent learns to solve a variety of kitchen or sorting tasks using a simulated robotic arm, as well as achieve arbitrary poses with a biped or a quadruped agent without any reward or demonstration supervision. I will further discuss how we can scale world model training by considering datasets of passive videos and other improvements to world models and planning. An agent that possesses a world model can use it to learn general knowledge from diverse datasets and to solve diverse tasks; I believe world models provide a principled and promising path towards building more and more general-purpose machines."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy (PhD)"]},{"key":"dc:title","label":"Title","values":["LEARNING TO ACT FROM DIVERSE DATA SOURCES VIA WORLD MODELS"]}]}],"canonical_facts":{"dc:contributor.advisor":["Daniilidis, Kostas","Levine, Sergey"],"dc:creator":["Rybkin, Oleh"],"dc:date.accessioned":["2023-11-22T15:55:51Z"],"dc:date.available":["2023-11-22T15:55:51Z"],"dc:date.issued":["2023"],"dc:description.abstract":["The next frontier in learning to act is generalization - the ability of the agent to operate in a diverse set of environments and to solve a diverse set of tasks. How can we learn generalist agents? I argue that learning world models that predict future outcomes of actions directly from image observations is a uniquely suitable approach for training generalizable agents. I will discuss how world models can enable powerful unsupervised exploration and how to use a single world model to learn a diverse range of tasks. An implementation of this agent learns to solve a variety of kitchen or sorting tasks using a simulated robotic arm, as well as achieve arbitrary poses with a biped or a quadruped agent without any reward or demonstration supervision. I will further discuss how we can scale world model training by considering datasets of passive videos and other improvements to world models and planning. An agent that possesses a world model can use it to learn general knowledge from diverse datasets and to solve diverse tasks; I believe world models provide a principled and promising path towards building more and more general-purpose machines."],"dc:description.degree":["Doctor of Philosophy (PhD)"],"dc:identifier.uri":["https://repository.upenn.edu/handle/20.500.14332/58993"],"dc:language.iso":["en"],"dc:subject":["Computer Sciences"],"dc:title":["LEARNING TO ACT FROM DIVERSE DATA SOURCES VIA WORLD MODELS"],"dc:type":["Dissertation/Thesis"]},"updated_at":"2026-07-24T03:45:26Z"}