{"id":{"repo_id":"usm","oai_identifier":"oai:aquila.usm.edu:masters_theses-1676"},"canonical_url":"https://search.dev.ndltd.org/etd/usm/oai:aquila.usm.edu:masters_theses-1676","repository":{"repo_id":"usm","name":"University of Southern Mississippi","base_url":"https://aquila.usm.edu/do/oai/"},"display":{"title":"Human Agent Transfer from Observations","abstract":"<p>Learning from human demonstration (LfD), among many speedup techniques for reinforcement learning (RL), has seen many successful applications. We consider one LfD technique called Human Agent Transfer (HAT), where a model of the human demonstrator’s decision function is induced via supervised learning, and used as an initial bias for RL. Some recent work in LfD have investigated learning from observations only, i.e., when only the demonstrator’s states (and not its actions) are available to the learner. Since the demonstrator’s actions are treated as labels for HAT, supervised learning becomes untenable in their absence. We adapt the idea of learning an inverse dynamics model from the data acquired by the learner’s interactions with the environment, and deploy it to fill in the missing actions of the demonstrator. The resulting version of HAT—called State-only HAT (SoHAT)—is experimentally shown to preserve some advantages of HAT in benchmark domains with both discrete and continuous actions. This thesis also establishes principled modifications of an existing baseline algorithm—called A3C—to create its HAT and SoHAT variants that are used in our experiments.</p>","abstract_html":"&lt;p&gt;Learning from human demonstration (LfD), among many speedup techniques for reinforcement learning (RL), has seen many successful applications. We consider one LfD technique called Human Agent Transfer (HAT), where a model of the human demonstrator’s decision function is induced via supervised learning, and used as an initial bias for RL. Some recent work in LfD have investigated learning from observations only, i.e., when only the demonstrator’s states (and not its actions) are available to the learner. Since the demonstrator’s actions are treated as labels for HAT, supervised learning becomes untenable in their absence. We adapt the idea of learning an inverse dynamics model from the data acquired by the learner’s interactions with the environment, and deploy it to fill in the missing actions of the demonstrator. The resulting version of HAT—called State-only HAT (SoHAT)—is experimentally shown to preserve some advantages of HAT in benchmark domains with both discrete and continuous actions. This thesis also establishes principled modifications of an existing baseline algorithm—called A3C—to create its HAT and SoHAT variants that are used in our experiments.&lt;/p&gt;","abstract_has_math":false,"creators":["Racharla, Sneha"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Bikramjit Banerjee","Dia Ali","Beddhu Murali"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-05-01T07:00:00Z","date_published":"2019-05-01T07:00:00Z","updated_at":"2026-07-24T05:45:12Z","subjects":["Reinforcement Learning","Human Agent Transfer","State-only Human Agent Transfer","Computational Engineering","Robotics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://aquila.usm.edu/masters_theses/630","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Bikramjit Banerjee","Dia Ali","Beddhu Murali"]},{"key":"dc:creator","label":"Author","values":["Racharla, Sneha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2029-05-09T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Human Agent Transfer","State-only Human Agent Transfer","Computational Engineering","Robotics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://aquila.usm.edu/masters_theses/630"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Learning from human demonstration (LfD), among many speedup techniques for reinforcement learning (RL), has seen many successful applications. We consider one LfD technique called Human Agent Transfer (HAT), where a model of the human demonstrator’s decision function is induced via supervised learning, and used as an initial bias for RL. Some recent work in LfD have investigated learning from observations only, i.e., when only the demonstrator’s states (and not its actions) are available to the learner. Since the demonstrator’s actions are treated as labels for HAT, supervised learning becomes untenable in their absence. We adapt the idea of learning an inverse dynamics model from the data acquired by the learner’s interactions with the environment, and deploy it to fill in the missing actions of the demonstrator. The resulting version of HAT—called State-only HAT (SoHAT)—is experimentally shown to preserve some advantages of HAT in benchmark domains with both discrete and continuous actions. This thesis also establishes principled modifications of an existing baseline algorithm—called A3C—to create its HAT and SoHAT variants that are used in our experiments.</p>"]},{"key":"dc:title","label":"Title","values":["Human Agent Transfer from Observations"]}]}],"canonical_facts":{"dc:contributor":["Bikramjit Banerjee","Dia Ali","Beddhu Murali"],"dc:creator":["Racharla, Sneha"],"dc:date.available":["2029-05-09T07:00:00Z"],"dc:description.abstract":["<p>Learning from human demonstration (LfD), among many speedup techniques for reinforcement learning (RL), has seen many successful applications. We consider one LfD technique called Human Agent Transfer (HAT), where a model of the human demonstrator’s decision function is induced via supervised learning, and used as an initial bias for RL. Some recent work in LfD have investigated learning from observations only, i.e., when only the demonstrator’s states (and not its actions) are available to the learner. Since the demonstrator’s actions are treated as labels for HAT, supervised learning becomes untenable in their absence. We adapt the idea of learning an inverse dynamics model from the data acquired by the learner’s interactions with the environment, and deploy it to fill in the missing actions of the demonstrator. The resulting version of HAT—called State-only HAT (SoHAT)—is experimentally shown to preserve some advantages of HAT in benchmark domains with both discrete and continuous actions. This thesis also establishes principled modifications of an existing baseline algorithm—called A3C—to create its HAT and SoHAT variants that are used in our experiments.</p>"],"dc:identifier":["https://aquila.usm.edu/masters_theses/630"],"dc:subject":["Reinforcement Learning","Human Agent Transfer","State-only Human Agent Transfer","Computational Engineering","Robotics"],"dc:title":["Human Agent Transfer from Observations"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:45:12Z"}