Brigham Young University - Provo
Limitations and Extensions of the WoLF-PHC Algorithm
Abstract
dc:description.abstractPolicy Hill Climbing (PHC) is a reinforcement learning algorithm that extends Q-learning to learn probabilistic policies for multi-agent games. WoLF-PHC extends PHC with the "win or learn fast" principle. A proof that PHC will diverge in self-play when playing Shapley's game is given, and WoLF-PHC is shown empirically to diverge as well. Various WoLF-PHC based modifications were created, evaluated, and compared in an attempt to obtain convergence to the single shot Nash equilibrium when playing Shapley's game in self-play without using more information than WoLF-PHC uses. Partial Commitment WoLF-PHC (PCWoLF-PHC), which performs best on Shapley's game, is tested on other matrix games and shown to produce satisfactory results.
Degree
thesis:*- Name thesis:degree_name
- MS
- Grantor dc:publisher
- Brigham Young University - Provo
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Cook, Philip R.
Subjects
dc:subject × 8Rights
- Language dc:language
- English
Identifiers
dc:identifier.*- Repository record dc:identifier
- https://scholarsarchive.byu.edu/etd/1222
- OAI identifier oai:identifier
- oai:scholarsarchive.byu.edu:etd-2221