{"id":{"repo_id":"usm","oai_identifier":"oai:aquila.usm.edu:masters_theses-1674"},"canonical_url":"https://search.dev.ndltd.org/etd/usm/oai:aquila.usm.edu:masters_theses-1674","repository":{"repo_id":"usm","name":"University of Southern Mississippi","base_url":"https://aquila.usm.edu/do/oai/"},"display":{"title":"Team Learning from Human Demonstration with Coordination Confidence","abstract":"<p>Among an array of techniques proposed to speed-up reinforcement learning (RL), learn- ing from human demonstration has a proven record of success. A related technique, called Human Agent Transfer (HAT), and its confidence-based derivatives have been successfully applied to single agent RL. This paper investigates their application to collaborative multi- agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view), and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.</p>","abstract_html":"&lt;p&gt;Among an array of techniques proposed to speed-up reinforcement learning (RL), learn- ing from human demonstration has a proven record of success. A related technique, called Human Agent Transfer (HAT), and its confidence-based derivatives have been successfully applied to single agent RL. This paper investigates their application to collaborative multi- agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view), and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.&lt;/p&gt;","abstract_has_math":false,"creators":["Vittanala, Syamala Nanditha"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Bikramjit Banerjee","Beddhu Murali","Dia Ali"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-05-01T07:00:00Z","date_published":"2019-05-01T07:00:00Z","updated_at":"2026-07-24T05:45:12Z","subjects":["Reinforcement Learning","Multi-Agent Domains","Learner Bootstrapping Methods","Human Agent Transfer","Computational Engineering","Robotics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://aquila.usm.edu/masters_theses/629","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Bikramjit Banerjee","Beddhu Murali","Dia Ali"]},{"key":"dc:creator","label":"Author","values":["Vittanala, Syamala Nanditha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2029-05-10T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Multi-Agent Domains","Learner Bootstrapping Methods","Human Agent Transfer","Computational Engineering","Robotics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://aquila.usm.edu/masters_theses/629"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Among an array of techniques proposed to speed-up reinforcement learning (RL), learn- ing from human demonstration has a proven record of success. A related technique, called Human Agent Transfer (HAT), and its confidence-based derivatives have been successfully applied to single agent RL. This paper investigates their application to collaborative multi- agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view), and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.</p>"]},{"key":"dc:title","label":"Title","values":["Team Learning from Human Demonstration with Coordination Confidence"]}]}],"canonical_facts":{"dc:contributor":["Bikramjit Banerjee","Beddhu Murali","Dia Ali"],"dc:creator":["Vittanala, Syamala Nanditha"],"dc:date.available":["2029-05-10T07:00:00Z"],"dc:description.abstract":["<p>Among an array of techniques proposed to speed-up reinforcement learning (RL), learn- ing from human demonstration has a proven record of success. A related technique, called Human Agent Transfer (HAT), and its confidence-based derivatives have been successfully applied to single agent RL. This paper investigates their application to collaborative multi- agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view), and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.</p>"],"dc:identifier":["https://aquila.usm.edu/masters_theses/629"],"dc:subject":["Reinforcement Learning","Multi-Agent Domains","Learner Bootstrapping Methods","Human Agent Transfer","Computational Engineering","Robotics"],"dc:title":["Team Learning from Human Demonstration with Coordination Confidence"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:45:12Z"}