{"id":{"repo_id":"uoit","oai_identifier":"oai:ontariotechu.scholaris.ca:10155/1706"},"canonical_url":"https://search.dev.ndltd.org/etd/uoit/oai:ontariotechu.scholaris.ca:10155/1706","repository":{"repo_id":"uoit","name":"Ontario Institute of Technology","base_url":"https://ontariotechu.scholaris.ca/server/oai/request"},"display":{"title":"Preferential proximal policy optimization in reinforcement learning","abstract":"The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, this Thesis introduces Preferential Proximal Policy Optimization (P3O), which integrates the importance of these states into parameter updates. We determine state importance by multiplying the variance of action probabilities by the value function, then normalizing and smoothing this with the Exponentially Weighted Moving Average (EWMA). This calculated importance is incorporated into the surrogate objective function, redefining value and advantage estimation in PPO. Our method auto-selects state importance, which can apply to any on-policy reinforcement learning algorithm using a value function. Empirical evaluations across six Atari environments demonstrate that our approach outperforms the baseline (vanilla PPO) across different tested environments, highlighting the value of our proposed method in learning complex environments.","abstract_html":"The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, this Thesis introduces Preferential Proximal Policy Optimization (P3O), which integrates the importance of these states into parameter updates. We determine state importance by multiplying the variance of action probabilities by the value function, then normalizing and smoothing this with the Exponentially Weighted Moving Average (EWMA). This calculated importance is incorporated into the surrogate objective function, redefining value and advantage estimation in PPO. Our method auto-selects state importance, which can apply to any on-policy reinforcement learning algorithm using a value function. Empirical evaluations across six Atari environments demonstrate that our approach outperforms the baseline (vanilla PPO) across different tested environments, highlighting the value of our proposed method in learning complex environments.","abstract_has_math":false,"creators":["Balasuntharam, Tamilselvan"],"institution":"University of Ontario Institute of Technology","degree_name":"Master of Science (MSc)","degree_level":null,"degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["Davoudi, Kourosh","Ebrahimi, Mehran"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12-01","date_published":"2023-12-01","updated_at":"2026-07-24T05:35:16Z","subjects":["Reinforcement learning","Policy gradient methods","Deep learning","Policy optimization"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10155/1706","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Davoudi, Kourosh","Ebrahimi, Mehran"]},{"key":"dc:creator","label":"Author","values":["Balasuntharam, Tamilselvan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-12-18T19:28:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-12-18T19:28:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-12-01"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Ontario Institute of Technology"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement learning","Policy gradient methods","Deep learning","Policy optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10155/1706"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, this Thesis introduces Preferential Proximal Policy Optimization (P3O), which integrates the importance of these states into parameter updates. We determine state importance by multiplying the variance of action probabilities by the value function, then normalizing and smoothing this with the Exponentially Weighted Moving Average (EWMA). This calculated importance is incorporated into the surrogate objective function, redefining value and advantage estimation in PPO. Our method auto-selects state importance, which can apply to any on-policy reinforcement learning algorithm using a value function. Empirical evaluations across six Atari environments demonstrate that our approach outperforms the baseline (vanilla PPO) across different tested environments, highlighting the value of our proposed method in learning complex environments."]},{"key":"dc:title","label":"Title","values":["Preferential proximal policy optimization in reinforcement learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Davoudi, Kourosh","Ebrahimi, Mehran"],"dc:creator":["Balasuntharam, Tamilselvan"],"dc:date.accessioned":["2023-12-18T19:28:29Z"],"dc:date.available":["2023-12-18T19:28:29Z"],"dc:date.issued":["2023-12-01"],"dc:description.abstract":["The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, this Thesis introduces Preferential Proximal Policy Optimization (P3O), which integrates the importance of these states into parameter updates. We determine state importance by multiplying the variance of action probabilities by the value function, then normalizing and smoothing this with the Exponentially Weighted Moving Average (EWMA). This calculated importance is incorporated into the surrogate objective function, redefining value and advantage estimation in PPO. Our method auto-selects state importance, which can apply to any on-policy reinforcement learning algorithm using a value function. Empirical evaluations across six Atari environments demonstrate that our approach outperforms the baseline (vanilla PPO) across different tested environments, highlighting the value of our proposed method in learning complex environments."],"dc:identifier.uri":["https://hdl.handle.net/10155/1706"],"dc:language.iso":["en"],"dc:subject":["Reinforcement learning","Policy gradient methods","Deep learning","Policy optimization"],"dc:title":["Preferential proximal policy optimization in reinforcement learning"],"dc:type":["Thesis"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["University of Ontario Institute of Technology"]},"updated_at":"2026-07-24T05:35:16Z"}