Rice University
Multilevel-in-Time Methods for Optimal Control of PDEs and Training of Recurrent Neural Networks
Abstract
dc:description.abstractThis thesis develops and analyzes “time”-parallel methods for two problem classes: Optimality systems arising in discretized linear-quadratic Partial Differential Equation (PDE)-constrained optimization problems and Recurrent Neural Network (RNN) training problems. While these two problem classes seem different, they share a time-like structure. The proposed methods exploit decompositions in time of the optimization problem or of some of its components to introduce parallelism into the solution process, to reduce the global memory requirements of the solution algorithm, or to mitigate numerical issues due to stability of the underlying equations. For the solution of PDE-constrained problems, the proposed approach builds on a time-domain decomposition of the optimality system into coupled time-subdomain optimality systems and an extension of the Parareal algorithm to this system. The resulting two-grid iteration is interpreted as a Schur complement iteration. This allows its extension to a multigrid iteration and it also reveals how to derive a preconditioner for use in a Krylov subspace methods. The performance of the multigrid iterations as well as of flexible GMRES with the multigrid preconditioner are illustrated for an optimal control problem governed by an advection-diffusion PDE. The optimization problem arising in the training of RNNs is interpreted as a discrete-time optimal control problem. This interpretation is used to extend methods originally designed for optimal control of (partial) differential equations to the training of RNNs. Specifically, this thesis formulates a Gauss-Newton-type method and shows that the computation of the Gauss-Newton step can be formulated as the solution of a discrete-time linear quadratic optimal control problem. This formulation provides a connection with the multigrid-in-time methods developed for PDE-constrained problems. This thesis also develops a convergence analysis for inexact stochastic gradient methods. These methods allow the use of small, randomly selected batches of training data during the optimization, and they allow biases in the gradient approximations. The analysis does not require convexity of the problem and extends recent convergence results for a biased stochastic first-order (SFO) method developed for (strongly) convex PDE-constrained optimization problems under uncertainty. Moreover, the convergence analysis developed in this thesis is applied to the recently proposed Multigrid Reduction in Time (MGRIT) training of implicit Gated Recurrent Units.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Engineering
- Grantor
- Rice University
- Year dc:date.issued
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lin, Shengchao
- Advisor dc:contributor.advisor
-
- Heinkenschloss, Matthias
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright is held by the author, unless otherwise indicated. Permission to reuse, publish, or reproduce the work beyond the bounds of fair use or other exemptions to copyright law must be obtained from the copyright holder.
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1911/113496
- OAI identifier oai:identifier
- oai:repository.rice.edu:1911/113496