Back to results

University of Ontario Institute of Technology

Preferential proximal policy optimization in reinforcement learning

Abstract

dc:description.abstract

The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, this Thesis introduces Preferential Proximal Policy Optimization (P3O), which integrates the importance of these states into parameter updates. We determine state importance by multiplying the variance of action probabilities by the value function, then normalizing and smoothing this with the Exponentially Weighted Moving Average (EWMA). This calculated importance is incorporated into the surrogate objective function, redefining value and advantage estimation in PPO. Our method auto-selects state importance, which can apply to any on-policy reinforcement learning algorithm using a value function. Empirical evaluations across six Atari environments demonstrate that our approach outperforms the baseline (vanilla PPO) across different tested environments, highlighting the value of our proposed method in learning complex environments.

Degree

thesis:*
Name thesis:degree_name
Master of Science (MSc)
Discipline thesis:degree_discipline
Electrical and Computer Engineering
Grantor
University of Ontario Institute of Technology
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Balasuntharam, Tamilselvan
Advisors dc:contributor.advisor
  • Davoudi, Kourosh
  • Ebrahimi, Mehran

Subjects

dc:subject × 4

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10155/1706
OAI identifier oai:identifier
oai:ontariotechu.scholaris.ca:10155/1706

Chain of custody

source
Harvested from
Ontario Institute of Technology
Base URL
ontariotechu.scholaris.ca/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Balasuntharam, Tamilselvan. Preferential proximal policy optimization in reinforcement learning. University of Ontario Institute of Technology, 2023. https://hdl.handle.net/10155/1706