Back to results

Virginia Tech

Sample Complexity of Incremental Policy Gradient Methods for Solving Multi-Task Reinforcement Learning

Abstract

dc:description.abstract

We consider a multi-task learning problem, where an agent is presented a number of N reinforcement learning tasks. To solve this problem, we are interested in studying the gradient approach, which iteratively updates an estimate of the optimal policy using the gradients of the value functions. The classic policy gradient method, however, may be expensive to implement in the multi-task settings as it requires access to the gradients of all the tasks at every iteration. To circumvent this issue, in this paper we propose to study an incremental policy gradient method, where the agent only uses the gradient of only one task at each iteration. Our main contribution is to provide theoretical results to characterize the performance of the proposed method. In particular, we show that incremental policy gradient methods converge to the optimal value of the multi-task reinforcement learning objectives at a sublinear rate O(1/√k), where k is the number of iterations. To illustrate its performance, we apply the proposed method to solve a simple multi-task variant of GridWorld problems, where an agent seeks to find an policy to navigate effectively in different environments.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Discipline thesis:degree_discipline
Electrical Engineering
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bai, Yitao
Chair dc:contributor.committeechair
  • Doan, Thinh T.
Committee members dc:contributor.committeemember
  • Stilwell, Daniel J.
  • Jin, Ming

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • CC0 1.0 Universal
Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10919/118699
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/118699

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Bai, Yitao. Sample Complexity of Incremental Policy Gradient Methods for Solving Multi-Task Reinforcement Learning. Virginia Tech, 2024. https://hdl.handle.net/10919/118699