{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/32995130"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/32995130","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Subdominance Minimization: A Satisficing Perspective on Imitation Learning","abstract":"Human decision-making often hinges on a delicately-balanced interplay between multiple, potentially conflicting objectives. However, prevailing imitation learning methods tend to prioritize optimizing a single imitation objective. This myopic focus on a singular objective frequently leads to unintended and undesirable behaviors in learned models. For example, an autonomous vehicle prioritizing travel time over adherence to traffic laws, or indeed a language model learning to generate increasingly creative responses at the expense of factual accuracy. We instead adopt the recently-proposed notion of subdominance which, given an arbitrary number of cost objectives to minimize simultaneously, quantifies the cost of choosing one solution relative to a reference solution. We present a novel focused satisficing approach to imitation learning, using subdominance to seek imitator policies that the demonstrator may consider acceptable, rather than optimal. We leverage both offline and online subdominance minimization to focus policy learning on parts of the demonstrator's trajectories which are hardest to imitate. In the domain of offline reinforcement learning, we present another approach for training decision transformers offline via subdominance minimization which allows us to learn autoregressive policies even in the absence of a ground truth reward function. Finally, we present an approach for language-conditioned, multitask imitation learning, where we leverage learned instruction representations to train language-conditioned subdominance minimizing policies.","abstract_html":"Human decision-making often hinges on a delicately-balanced interplay between multiple, potentially conflicting objectives. However, prevailing imitation learning methods tend to prioritize optimizing a single imitation objective. This myopic focus on a singular objective frequently leads to unintended and undesirable behaviors in learned models. For example, an autonomous vehicle prioritizing travel time over adherence to traffic laws, or indeed a language model learning to generate increasingly creative responses at the expense of factual accuracy. We instead adopt the recently-proposed notion of subdominance which, given an arbitrary number of cost objectives to minimize simultaneously, quantifies the cost of choosing one solution relative to a reference solution. We present a novel focused satisficing approach to imitation learning, using subdominance to seek imitator policies that the demonstrator may consider acceptable, rather than optimal. We leverage both offline and online subdominance minimization to focus policy learning on parts of the demonstrator&#x27;s trajectories which are hardest to imitate. In the domain of offline reinforcement learning, we present another approach for training decision transformers offline via subdominance minimization which allows us to learn autoregressive policies even in the absence of a ground truth reward function. Finally, we present an approach for language-conditioned, multitask imitation learning, where we leverage learned instruction representations to train language-conditioned subdominance minimizing policies.","abstract_has_math":false,"creators":["Rushit Shah (24400088)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-05-01T00:00:00Z","date_published":"2026-05-01T00:00:00Z","updated_at":"2026-07-27T21:33:48Z","subjects":["Artificial Intelligence","Machine Learning","Reinforcement Learning","Imitation Learning","Decision Theory","Behavioral Economics"],"languages":[],"rights":["In Copyright","Open Access after 2028-05-01"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.32995130.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Rushit Shah (24400088)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-05-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Subdominance_Minimization_A_Satisficing_Perspective_on_Imitation_Learning/32995130"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial Intelligence","Machine Learning","Reinforcement Learning","Imitation Learning","Decision Theory","Behavioral Economics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright","Open Access after 2028-05-01"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.32995130.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Human decision-making often hinges on a delicately-balanced interplay between multiple, potentially conflicting objectives. However, prevailing imitation learning methods tend to prioritize optimizing a single imitation objective. This myopic focus on a singular objective frequently leads to unintended and undesirable behaviors in learned models. For example, an autonomous vehicle prioritizing travel time over adherence to traffic laws, or indeed a language model learning to generate increasingly creative responses at the expense of factual accuracy. We instead adopt the recently-proposed notion of subdominance which, given an arbitrary number of cost objectives to minimize simultaneously, quantifies the cost of choosing one solution relative to a reference solution. We present a novel focused satisficing approach to imitation learning, using subdominance to seek imitator policies that the demonstrator may consider acceptable, rather than optimal. We leverage both offline and online subdominance minimization to focus policy learning on parts of the demonstrator's trajectories which are hardest to imitate. In the domain of offline reinforcement learning, we present another approach for training decision transformers offline via subdominance minimization which allows us to learn autoregressive policies even in the absence of a ground truth reward function. Finally, we present an approach for language-conditioned, multitask imitation learning, where we leverage learned instruction representations to train language-conditioned subdominance minimizing policies."]},{"key":"dc:title","label":"Title","values":["Subdominance Minimization: A Satisficing Perspective on Imitation Learning"]}]}],"canonical_facts":{"dc:creator":["Rushit Shah (24400088)"],"dc:date":["2026-05-01T00:00:00Z"],"dc:description":["Human decision-making often hinges on a delicately-balanced interplay between multiple, potentially conflicting objectives. However, prevailing imitation learning methods tend to prioritize optimizing a single imitation objective. This myopic focus on a singular objective frequently leads to unintended and undesirable behaviors in learned models. For example, an autonomous vehicle prioritizing travel time over adherence to traffic laws, or indeed a language model learning to generate increasingly creative responses at the expense of factual accuracy. We instead adopt the recently-proposed notion of subdominance which, given an arbitrary number of cost objectives to minimize simultaneously, quantifies the cost of choosing one solution relative to a reference solution. We present a novel focused satisficing approach to imitation learning, using subdominance to seek imitator policies that the demonstrator may consider acceptable, rather than optimal. We leverage both offline and online subdominance minimization to focus policy learning on parts of the demonstrator's trajectories which are hardest to imitate. In the domain of offline reinforcement learning, we present another approach for training decision transformers offline via subdominance minimization which allows us to learn autoregressive policies even in the absence of a ground truth reward function. Finally, we present an approach for language-conditioned, multitask imitation learning, where we leverage learned instruction representations to train language-conditioned subdominance minimizing policies."],"dc:identifier":["10.25417/uic.32995130.v1"],"dc:relation":["https://figshare.com/articles/thesis/Subdominance_Minimization_A_Satisficing_Perspective_on_Imitation_Learning/32995130"],"dc:rights":["In Copyright","Open Access after 2028-05-01"],"dc:subject":["Artificial Intelligence","Machine Learning","Reinforcement Learning","Imitation Learning","Decision Theory","Behavioral Economics"],"dc:title":["Subdominance Minimization: A Satisficing Perspective on Imitation Learning"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:33:48Z"}