{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105916"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105916","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large-scale training of deep neural networks","abstract":"Accelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.","abstract_html":"Accelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.","abstract_has_math":false,"creators":["Dryden, Nikoli Joseph"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Snir, Marc","Gropp, William","Hwu, Wen-mei","Van Essen, Brian","Schwing, Alexander"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:59:36Z","date_published":"2019-11-26T20:59:36Z","updated_at":"2026-07-22T22:24:45Z","subjects":["High-performance computing","deep learning","convolutional neural network","parallel computing","machine learning"],"languages":["en"],"rights":["Copyright 2019 Nikoli Joseph Dryden"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105916","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Snir, Marc","Gropp, William","Hwu, Wen-mei","Van Essen, Brian","Schwing, Alexander"]},{"key":"dc:creator","label":"Author","values":["Dryden, Nikoli Joseph"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:59:36Z","2021-11-27T10:15:23Z","2019-07-09","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["High-performance computing","deep learning","convolutional neural network","parallel computing","machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Nikoli Joseph Dryden"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105916"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Accelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-08-01","The student, Nikoli Dryden, accepted the attached license on 2019-07-09 at 14:44.","The student, Nikoli Dryden, submitted this Dissertation for approval on 2019-07-09 at 14:44.","This Dissertation was approved for publication on 2019-07-09 at 15:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14222 on 2019-11-26 at 14:01:58","Made available in DSpace on 2019-11-26T20:59:36Z (GMT). No. of bitstreams: 2 DRYDEN-DISSERTATION-2019.pdf: 1517701 bytes, checksum: e62a499063edb4ed5333cccdef479754 (MD5) LICENSE.txt: 4210 bytes, checksum: 62bebd64df1dccf769c6f4d6b266b4b5 (MD5) Previous issue date: 2019-07-09","Embargo set by: Seth Robbins for item 113063 Lift date: 2021-11-26T20:59:54Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113063 on 2021-11-27T10:15:23Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large-scale training of deep neural networks"]}]}],"canonical_facts":{"dc:contributor":["Snir, Marc","Gropp, William","Hwu, Wen-mei","Van Essen, Brian","Schwing, Alexander"],"dc:creator":["Dryden, Nikoli Joseph"],"dc:date":["2019-11-26T20:59:36Z","2021-11-27T10:15:23Z","2019-07-09","2019-08"],"dc:description":["Accelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-08-01","The student, Nikoli Dryden, accepted the attached license on 2019-07-09 at 14:44.","The student, Nikoli Dryden, submitted this Dissertation for approval on 2019-07-09 at 14:44.","This Dissertation was approved for publication on 2019-07-09 at 15:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14222 on 2019-11-26 at 14:01:58","Made available in DSpace on 2019-11-26T20:59:36Z (GMT). No. of bitstreams: 2 DRYDEN-DISSERTATION-2019.pdf: 1517701 bytes, checksum: e62a499063edb4ed5333cccdef479754 (MD5) LICENSE.txt: 4210 bytes, checksum: 62bebd64df1dccf769c6f4d6b266b4b5 (MD5) Previous issue date: 2019-07-09","Embargo set by: Seth Robbins for item 113063 Lift date: 2021-11-26T20:59:54Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113063 on 2021-11-27T10:15:23Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105916"],"dc:language":["en"],"dc:rights":["Copyright 2019 Nikoli Joseph Dryden"],"dc:subject":["High-performance computing","deep learning","convolutional neural network","parallel computing","machine learning"],"dc:title":["Large-scale training of deep neural networks"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}