{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106434"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106434","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Accelerating distributed neural network training with network-centric approach","abstract":"Distributed training of Deep Neural Networks (DNN) is an important technique to reduce the training time of large DNNs for a wide range of applications. In existing distributed training approaches, however, the communication time to periodically exchange parameters (i.e., weights) and gradients among computer nodes over the network constitutes a large fraction of the total training time. To reduce the communication time, we propose an algorithm/hardware co-design, INCEPTIONN. More specifically, observing that gradients are much more tolerant to precision loss than parameters, we first propose a gradient-centric distributed training algorithm. As designed to exchange only gradients among nodes in a distributed manner, it can transfer less information, better overlap communication with computation, and apply a more aggressive lossy compression algorithm to all the information exchanged among nodes than traditional distributed algorithms. Second, exploiting unique characteristics of gradient values, we propose a lossy compression algorithm, optimized for compressing gradients. It accomplishes high compression ratios for compressing gradients without notably affecting the accuracy of trained DNNs. Lastly, we demonstrate that compression algorithms consume a large amount of CPU time, which in turn increases total training time albeit reduced communication time. To tackle this, we propose an in-network computing approach that delegates the lossy compression task to hardware integrated with a Network Interface Card (NIC). Our experiments show that INCEPTIONN can reduce a large portion of the communication time and thus the training time of DNNs, with little degradation in accuracy of trained DNNs.","abstract_html":"Distributed training of Deep Neural Networks (DNN) is an important technique to reduce the training time of large DNNs for a wide range of applications. In existing distributed training approaches, however, the communication time to periodically exchange parameters (i.e., weights) and gradients among computer nodes over the network constitutes a large fraction of the total training time. To reduce the communication time, we propose an algorithm/hardware co-design, INCEPTIONN. More specifically, observing that gradients are much more tolerant to precision loss than parameters, we first propose a gradient-centric distributed training algorithm. As designed to exchange only gradients among nodes in a distributed manner, it can transfer less information, better overlap communication with computation, and apply a more aggressive lossy compression algorithm to all the information exchanged among nodes than traditional distributed algorithms. Second, exploiting unique characteristics of gradient values, we propose a lossy compression algorithm, optimized for compressing gradients. It accomplishes high compression ratios for compressing gradients without notably affecting the accuracy of trained DNNs. Lastly, we demonstrate that compression algorithms consume a large amount of CPU time, which in turn increases total training time albeit reduced communication time. To tackle this, we propose an in-network computing approach that delegates the lossy compression task to hardware integrated with a Network Interface Card (NIC). Our experiments show that INCEPTIONN can reduce a large portion of the communication time and thus the training time of DNNs, with little degradation in accuracy of trained DNNs.","abstract_has_math":false,"creators":["Yuan, Yifan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Kim, Nam Sung"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T22:38:39Z","date_published":"2020-03-02T22:38:39Z","updated_at":"2026-07-22T22:24:47Z","subjects":["distributed training","accelerator"],"languages":["en"],"rights":["Copyright 2019 Yifan Yuan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106434","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kim, Nam Sung"]},{"key":"dc:creator","label":"Author","values":["Yuan, Yifan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T22:38:39Z","2022-03-03T10:15:22Z","2019-10-17","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["distributed training","accelerator"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yifan Yuan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106434"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Distributed training of Deep Neural Networks (DNN) is an important technique to reduce the training time of large DNNs for a wide range of applications. In existing distributed training approaches, however, the communication time to periodically exchange parameters (i.e., weights) and gradients among computer nodes over the network constitutes a large fraction of the total training time. To reduce the communication time, we propose an algorithm/hardware co-design, INCEPTIONN. More specifically, observing that gradients are much more tolerant to precision loss than parameters, we first propose a gradient-centric distributed training algorithm. As designed to exchange only gradients among nodes in a distributed manner, it can transfer less information, better overlap communication with computation, and apply a more aggressive lossy compression algorithm to all the information exchanged among nodes than traditional distributed algorithms. Second, exploiting unique characteristics of gradient values, we propose a lossy compression algorithm, optimized for compressing gradients. It accomplishes high compression ratios for compressing gradients without notably affecting the accuracy of trained DNNs. Lastly, we demonstrate that compression algorithms consume a large amount of CPU time, which in turn increases total training time albeit reduced communication time. To tackle this, we propose an in-network computing approach that delegates the lossy compression task to hardware integrated with a Network Interface Card (NIC). Our experiments show that INCEPTIONN can reduce a large portion of the communication time and thus the training time of DNNs, with little degradation in accuracy of trained DNNs.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Yifan Yuan, accepted the attached license on 2019-10-16 at 11:28.","The student, Yifan Yuan, submitted this Thesis for approval on 2019-10-17 at 11:04.","This Thesis was approved for publication on 2019-10-17 at 14:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14499 on 2020-02-28 at 17:35:46","Made available in DSpace on 2020-03-02T22:38:39Z (GMT). No. of bitstreams: 2 YUAN-THESIS-2019.pdf: 697509 bytes, checksum: 04c13b288846b999c63ebe53c4453fe7 (MD5) LICENSE.txt: 4207 bytes, checksum: 2f7056da77d43a35b3abbba998cb115c (MD5) Previous issue date: 2019-10-17","Embargo set by: Seth Robbins for item 113978 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113978 on 2022-03-03T10:15:22Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Accelerating distributed neural network training with network-centric approach"]}]}],"canonical_facts":{"dc:contributor":["Kim, Nam Sung"],"dc:creator":["Yuan, Yifan"],"dc:date":["2020-03-02T22:38:39Z","2022-03-03T10:15:22Z","2019-10-17","2019-12"],"dc:description":["Distributed training of Deep Neural Networks (DNN) is an important technique to reduce the training time of large DNNs for a wide range of applications. In existing distributed training approaches, however, the communication time to periodically exchange parameters (i.e., weights) and gradients among computer nodes over the network constitutes a large fraction of the total training time. To reduce the communication time, we propose an algorithm/hardware co-design, INCEPTIONN. More specifically, observing that gradients are much more tolerant to precision loss than parameters, we first propose a gradient-centric distributed training algorithm. As designed to exchange only gradients among nodes in a distributed manner, it can transfer less information, better overlap communication with computation, and apply a more aggressive lossy compression algorithm to all the information exchanged among nodes than traditional distributed algorithms. Second, exploiting unique characteristics of gradient values, we propose a lossy compression algorithm, optimized for compressing gradients. It accomplishes high compression ratios for compressing gradients without notably affecting the accuracy of trained DNNs. Lastly, we demonstrate that compression algorithms consume a large amount of CPU time, which in turn increases total training time albeit reduced communication time. To tackle this, we propose an in-network computing approach that delegates the lossy compression task to hardware integrated with a Network Interface Card (NIC). Our experiments show that INCEPTIONN can reduce a large portion of the communication time and thus the training time of DNNs, with little degradation in accuracy of trained DNNs.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Yifan Yuan, accepted the attached license on 2019-10-16 at 11:28.","The student, Yifan Yuan, submitted this Thesis for approval on 2019-10-17 at 11:04.","This Thesis was approved for publication on 2019-10-17 at 14:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14499 on 2020-02-28 at 17:35:46","Made available in DSpace on 2020-03-02T22:38:39Z (GMT). No. of bitstreams: 2 YUAN-THESIS-2019.pdf: 697509 bytes, checksum: 04c13b288846b999c63ebe53c4453fe7 (MD5) LICENSE.txt: 4207 bytes, checksum: 2f7056da77d43a35b3abbba998cb115c (MD5) Previous issue date: 2019-10-17","Embargo set by: Seth Robbins for item 113978 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113978 on 2022-03-03T10:15:22Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106434"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yifan Yuan"],"dc:subject":["distributed training","accelerator"],"dc:title":["Accelerating distributed neural network training with network-centric approach"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}