{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/36116"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/36116","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Learning Strategies in Multi-Agent Systems - Applications to the Herding Problem","abstract":"\"Multi-Agent systems\" is a topic for a lot of research, especially research involving strategy, evolution and cooperation among various agents. Various learning algorithm schemes have been proposed such as reinforcement learning and evolutionary computing. In this thesis two solutions to a multi-agent herding problem are presented. One solution is based on Q-learning algorithm, while the other is based on modeling of artificial immune system. Q-learning solution for the herding problem is developed, using region-based local learning for each individual agent. Individual and batch processing reinforcement algorithms are implemented for non-cooperative agents. Agents in this formulation do not share any information or knowledge. Issues such as computational requirements, and convergence are discussed. An idiotopic artificial immune network is proposed that includes individual B-cell model for agents and T-cell model for controlling the interaction among these agents. Two network models are proposed--one for evolving group behavior/strategy arbitration and the other for individual action selection. A comparative study of the Q-learning solution and the immune network solution is done on important aspects such as computation requirements, predictability, and convergence.","abstract_html":"&quot;Multi-Agent systems&quot; is a topic for a lot of research, especially research involving strategy, evolution and cooperation among various agents. Various learning algorithm schemes have been proposed such as reinforcement learning and evolutionary computing. In this thesis two solutions to a multi-agent herding problem are presented. One solution is based on Q-learning algorithm, while the other is based on modeling of artificial immune system. Q-learning solution for the herding problem is developed, using region-based local learning for each individual agent. Individual and batch processing reinforcement algorithms are implemented for non-cooperative agents. Agents in this formulation do not share any information or knowledge. Issues such as computational requirements, and convergence are discussed. An idiotopic artificial immune network is proposed that includes individual B-cell model for agents and T-cell model for controlling the interaction among these agents. Two network models are proposed--one for evolving group behavior/strategy arbitration and the other for individual action selection. A comparative study of the Q-learning solution and the immune network solution is done on important aspects such as computation requirements, predictability, and convergence.","abstract_has_math":false,"creators":["Gadre, Aditya Shrikant"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Electrical and Computer Engineering","degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Kachroo, Pushkin"],"committee_members":["VanLandingham, Hugh F.","Saunders, William R."],"year":2001,"date_issued":"2001-11-30","date_published":"2001-11-30","updated_at":"2026-07-22T22:19:47Z","subjects":["Idiotopic Network","Reinforcement Learning","Reward functions","Dynamic Programming","Q-learning","Artificial Immune System"],"languages":[],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["etd-12142001-002614"],"render_values":[{"text":"etd-12142001-002614","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10919/36116","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Kachroo, Pushkin"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["VanLandingham, Hugh F.","Saunders, William R."]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Gadre, Aditya Shrikant"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2014-03-14T20:49:30Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2014-03-14T20:49:30Z","2002-12-14"]},{"key":"dc:date.issued","label":"Date","values":["2001-11-30"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Idiotopic Network","Reinforcement Learning","Reward functions","Dynamic Programming","Q-learning","Artificial Immune System"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["etd-12142001-002614"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10919/36116"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["\"Multi-Agent systems\" is a topic for a lot of research, especially research involving strategy, evolution and cooperation among various agents. Various learning algorithm schemes have been proposed such as reinforcement learning and evolutionary computing. In this thesis two solutions to a multi-agent herding problem are presented. One solution is based on Q-learning algorithm, while the other is based on modeling of artificial immune system. Q-learning solution for the herding problem is developed, using region-based local learning for each individual agent. Individual and batch processing reinforcement algorithms are implemented for non-cooperative agents. Agents in this formulation do not share any information or knowledge. Issues such as computational requirements, and convergence are discussed. An idiotopic artificial immune network is proposed that includes individual B-cell model for agents and T-cell model for controlling the interaction among these agents. Two network models are proposed--one for evolving group behavior/strategy arbitration and the other for individual action selection. A comparative study of the Q-learning solution and the immune network solution is done on important aspects such as computation requirements, predictability, and convergence."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:title","label":"Title","values":["Learning Strategies in Multi-Agent Systems - Applications to the Herding Problem"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Kachroo, Pushkin"],"dc:contributor.committeemember":["VanLandingham, Hugh F.","Saunders, William R."],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Gadre, Aditya Shrikant"],"dc:date.accessioned":["2014-03-14T20:49:30Z"],"dc:date.available":["2014-03-14T20:49:30Z","2002-12-14"],"dc:date.issued":["2001-11-30"],"dc:description.abstract":["\"Multi-Agent systems\" is a topic for a lot of research, especially research involving strategy, evolution and cooperation among various agents. Various learning algorithm schemes have been proposed such as reinforcement learning and evolutionary computing. In this thesis two solutions to a multi-agent herding problem are presented. One solution is based on Q-learning algorithm, while the other is based on modeling of artificial immune system. Q-learning solution for the herding problem is developed, using region-based local learning for each individual agent. Individual and batch processing reinforcement algorithms are implemented for non-cooperative agents. Agents in this formulation do not share any information or knowledge. Issues such as computational requirements, and convergence are discussed. An idiotopic artificial immune network is proposed that includes individual B-cell model for agents and T-cell model for controlling the interaction among these agents. Two network models are proposed--one for evolving group behavior/strategy arbitration and the other for individual action selection. A comparative study of the Q-learning solution and the immune network solution is done on important aspects such as computation requirements, predictability, and convergence."],"dc:description.degree":["Master of Science"],"dc:identifier.other":["etd-12142001-002614"],"dc:identifier.uri":["http://hdl.handle.net/10919/36116"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Idiotopic Network","Reinforcement Learning","Reward functions","Dynamic Programming","Q-learning","Artificial Immune System"],"dc:title":["Learning Strategies in Multi-Agent Systems - Applications to the Herding Problem"],"dc:type":["Thesis"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:47Z"}