{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/139186"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/139186","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Robustness and Adaptation via a Generative Model of Policies in Reinforcement Learning","abstract":"In the natural world, life has found an uncountable number of ways to survive and often thrive. Between and even within species, each individual has a slightly unique way of existing, and this diversity lends robustness to life in general. In this work, we aim to incentivize diversity of agent policies while optimizing for an external reward. To this end, we introduce a generative model of policies which maps a low-dimensional latent space to an agent policy space. In order to learn a broad range of solutions, our generative model uses a diversity regularizer that incentivizes different agent behaviors given the same state. Agents are assigned a specific latent vector persistent throughout their trajectory, and the generator learns to encode behavioral preferences in the latent space. Results show that our generator is able to find an array of policies that can express agent individuality through distinct and unique agent policies. Of particular interest, we find that having a diverse policy space allows us to rapidly adapt to unforeseen environmental ablations simply by optimizing generated policies in the low-dimensional latent space. We test this adaptability in an open-ended grid-world, as well as in a competitive, zero-sum, two-player soccer environment.","abstract_html":"In the natural world, life has found an uncountable number of ways to survive and often thrive. Between and even within species, each individual has a slightly unique way of existing, and this diversity lends robustness to life in general. In this work, we aim to incentivize diversity of agent policies while optimizing for an external reward. To this end, we introduce a generative model of policies which maps a low-dimensional latent space to an agent policy space. In order to learn a broad range of solutions, our generative model uses a diversity regularizer that incentivizes different agent behaviors given the same state. Agents are assigned a specific latent vector persistent throughout their trajectory, and the generator learns to encode behavioral preferences in the latent space. Results show that our generator is able to find an array of policies that can express agent individuality through distinct and unique agent policies. Of particular interest, we find that having a diverse policy space allows us to rapidly adapt to unforeseen environmental ablations simply by optimizing generated policies in the low-dimensional latent space. We test this adaptability in an open-ended grid-world, as well as in a competitive, zero-sum, two-player soccer environment.","abstract_has_math":false,"creators":["Derek, Kenneth"],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science","school":null,"contributors":[],"advisors":["Isola, Phillip"],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-06","date_published":"2021-06","updated_at":"2026-07-22T22:21:40Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"rights_urls":["http://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/139186","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Isola, Phillip"]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Derek, Kenneth"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-01-14T14:55:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-01-14T14:55:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2021-06"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master","Master of Engineering in Electrical Engineering and Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright MIT"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/139186"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In the natural world, life has found an uncountable number of ways to survive and often thrive. Between and even within species, each individual has a slightly unique way of existing, and this diversity lends robustness to life in general. In this work, we aim to incentivize diversity of agent policies while optimizing for an external reward. To this end, we introduce a generative model of policies which maps a low-dimensional latent space to an agent policy space. In order to learn a broad range of solutions, our generative model uses a diversity regularizer that incentivizes different agent behaviors given the same state. Agents are assigned a specific latent vector persistent throughout their trajectory, and the generator learns to encode behavioral preferences in the latent space. Results show that our generator is able to find an array of policies that can express agent individuality through distinct and unique agent policies. Of particular interest, we find that having a diverse policy space allows us to rapidly adapt to unforeseen environmental ablations simply by optimizing generated policies in the low-dimensional latent space. We test this adaptability in an open-ended grid-world, as well as in a competitive, zero-sum, two-player soccer environment."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:title","label":"Title","values":["Robustness and Adaptation via a Generative Model of Policies in Reinforcement Learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Isola, Phillip"],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"],"dc:creator":["Derek, Kenneth"],"dc:date.accessioned":["2022-01-14T14:55:29Z"],"dc:date.available":["2022-01-14T14:55:29Z"],"dc:date.issued":["2021-06"],"dc:description.abstract":["In the natural world, life has found an uncountable number of ways to survive and often thrive. Between and even within species, each individual has a slightly unique way of existing, and this diversity lends robustness to life in general. In this work, we aim to incentivize diversity of agent policies while optimizing for an external reward. To this end, we introduce a generative model of policies which maps a low-dimensional latent space to an agent policy space. In order to learn a broad range of solutions, our generative model uses a diversity regularizer that incentivizes different agent behaviors given the same state. Agents are assigned a specific latent vector persistent throughout their trajectory, and the generator learns to encode behavioral preferences in the latent space. Results show that our generator is able to find an array of policies that can express agent individuality through distinct and unique agent policies. Of particular interest, we find that having a diverse policy space allows us to rapidly adapt to unforeseen environmental ablations simply by optimizing generated policies in the low-dimensional latent space. We test this adaptability in an open-ended grid-world, as well as in a competitive, zero-sum, two-player soccer environment."],"dc:description.degree":["M.Eng."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/139186"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"dc:rights.uri":["http://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Robustness and Adaptation via a Generative Model of Policies in Reinforcement Learning"],"dc:type":["Thesis"],"thesis:degree_name":["Master","Master of Engineering in Electrical Engineering and Computer Science"]},"updated_at":"2026-07-22T22:21:40Z"}