{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/117774"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/117774","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Communication-centric cross-stack acceleration for distributed machine learning","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_has_math":false,"creators":["Li, Youjie"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Kim, Nam Sung","Hwu, Wen-mei","Torrellas, Josep","Chen, Deming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12","date_published":"2022-12","updated_at":"2026-07-22T22:24:56Z","subjects":["Distributed Machine Learning","Distributed Training","Parallel Computing","In-network Computing","Smart Nics","Programmable Switches","Pipelined Sgd","Machine Learning Framework"],"languages":["en","eng"],"rights":["In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of [university/educational entity's name goes here]'s products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/117774","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kim, Nam Sung","Hwu, Wen-mei","Torrellas, Josep","Chen, Deming"]},{"key":"dc:creator","label":"Author","values":["Li, Youjie"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-12","2022-11-23"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Distributed Machine Learning","Distributed Training","Parallel Computing","In-network Computing","Smart Nics","Programmable Switches","Pipelined Sgd","Machine Learning Framework"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of [university/educational entity's name goes here]'s products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/117774"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Youjie Li, accepted the attached license on 2022-11-22 at 20:06.","The student, Youjie Li, submitted this Dissertation for approval on 2022-11-22 at 21:09.","This Dissertation was approved for publication on 2022-11-23 at 13:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18624 on 2023-04-12 at 07:31:54","Distributed training has been the Holy Grail in machine learning systems, as it is an indispensable technique for addressing the ever-growing computation demands and memory requirements due to the unprecedented scaling of model sizes and data volumes. However, even distributed training takes inordinate time, of which a large fraction is paid for communication overhead in either inter-server networks or intra-server interconnects. In this dissertation, we propose cross-stack solutions that span hardware, software, and algorithms to accelerate and scale distributed training systems. First, we propose in-network computing by leveraging novel network hardware in modern data-centers, such as programmable network interface cards and switches, for not only compressing traffic volume in real time but also reducing network hops on the fly. Second, we present new algorithms surrounding efficient communication, such as gradient compression and pipelined computation with communication, for further shrinking the network overhead while maintaining the training convergence and model accuracy. Third, we develop a next-generation machine learning framework by novel schemes of task decomposition and late binding, to train massive models even without sufficient memory while drastically reducing communication overhead within a server's interconnects."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Communication-centric cross-stack acceleration for distributed machine learning"]}]}],"canonical_facts":{"dc:contributor":["Kim, Nam Sung","Hwu, Wen-mei","Torrellas, Josep","Chen, Deming"],"dc:creator":["Li, Youjie"],"dc:date":["2022-12","2022-11-23"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Youjie Li, accepted the attached license on 2022-11-22 at 20:06.","The student, Youjie Li, submitted this Dissertation for approval on 2022-11-22 at 21:09.","This Dissertation was approved for publication on 2022-11-23 at 13:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18624 on 2023-04-12 at 07:31:54","Distributed training has been the Holy Grail in machine learning systems, as it is an indispensable technique for addressing the ever-growing computation demands and memory requirements due to the unprecedented scaling of model sizes and data volumes. However, even distributed training takes inordinate time, of which a large fraction is paid for communication overhead in either inter-server networks or intra-server interconnects. In this dissertation, we propose cross-stack solutions that span hardware, software, and algorithms to accelerate and scale distributed training systems. First, we propose in-network computing by leveraging novel network hardware in modern data-centers, such as programmable network interface cards and switches, for not only compressing traffic volume in real time but also reducing network hops on the fly. Second, we present new algorithms surrounding efficient communication, such as gradient compression and pipelined computation with communication, for further shrinking the network overhead while maintaining the training convergence and model accuracy. Third, we develop a next-generation machine learning framework by novel schemes of task decomposition and late binding, to train massive models even without sufficient memory while drastically reducing communication overhead within a server's interconnects."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/117774"],"dc:language":["en","eng"],"dc:rights":["In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of [university/educational entity's name goes here]'s products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink."],"dc:subject":["Distributed Machine Learning","Distributed Training","Parallel Computing","In-network Computing","Smart Nics","Programmable Switches","Pipelined Sgd","Machine Learning Framework"],"dc:title":["Communication-centric cross-stack acceleration for distributed machine learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}