University of Illinois at Urbana-Champaign
Large-scale training of deep neural networks
Abstract
dc:descriptionAccelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2019
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Dryden, Nikoli Joseph
- Contributors dc:contributor
-
- Snir, Marc
- Gropp, William
- Hwu, Wen-mei
- Van Essen, Brian
- Schwing, Alexander
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2019 Nikoli Joseph Dryden
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/105916
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/105916