{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/80294"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/80294","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Software-Hardware Optimizations for Efficient Collective Communications in Distributed Machine Learning Platforms","abstract":"Foundation machine learning (ML) models have emerged as one of the most prominent applications in modern computing, exemplified by mixture-of-experts–based large language models. The immense resource demands of these models have driven the development of large-scale, high-performance computing platforms tailored for artificial intelligence workloads. In such distributed platforms, both model parameters and data are partitioned and processed across numerous neural processing units, requiring frequent synchronization of activations and gradients through collective communication operations. As collective communication constitutes a primary bottleneck in distributed ML, optimizing its efficiency remains a critical research challenge. This dissertation explores software-hardware optimizations for collective communications to better understand the tightly coupled design space of networking in distributed ML platforms. First, it introduces ASTRA-sim2.0, an end-to-end simulation and modeling framework that enables comprehensive design space exploration of distributed ML platforms with arbitrary parallelization strategies and multi-dimensional networks. Second, it presents LIBRA, which enhances the bandwidth utilization of hierarchical collective communication algorithms by optimizing multi-dimensional network topologies via analytical modeling. Finally, the dissertation proposes two collective communication algorithm synthesizers, TACOS and PCCL, which automatically generate optimized collective communication algorithms for arbitrary network topologies through algorithmic approaches. Together, the dissertation underscores the significance of judicious software-hardware approaches in achieving efficient collective communication for large-scale distributed ML platforms.","abstract_html":"Foundation machine learning (ML) models have emerged as one of the most prominent applications in modern computing, exemplified by mixture-of-experts–based large language models. The immense resource demands of these models have driven the development of large-scale, high-performance computing platforms tailored for artificial intelligence workloads. In such distributed platforms, both model parameters and data are partitioned and processed across numerous neural processing units, requiring frequent synchronization of activations and gradients through collective communication operations. As collective communication constitutes a primary bottleneck in distributed ML, optimizing its efficiency remains a critical research challenge. This dissertation explores software-hardware optimizations for collective communications to better understand the tightly coupled design space of networking in distributed ML platforms. First, it introduces ASTRA-sim2.0, an end-to-end simulation and modeling framework that enables comprehensive design space exploration of distributed ML platforms with arbitrary parallelization strategies and multi-dimensional networks. Second, it presents LIBRA, which enhances the bandwidth utilization of hierarchical collective communication algorithms by optimizing multi-dimensional network topologies via analytical modeling. Finally, the dissertation proposes two collective communication algorithm synthesizers, TACOS and PCCL, which automatically generate optimized collective communication algorithms for arbitrary network topologies through algorithmic approaches. Together, the dissertation underscores the significance of judicious software-hardware approaches in achieving efficient collective communication for large-scale distributed ML platforms.","abstract_has_math":false,"creators":["Won, William Jonghoon"],"institution":"Georgia Institute of Technology","degree_name":"Computer Science, PhD","degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Krishna, Tushar"],"committee_chairs":[],"committee_members":["Mahajan, Divya","Lin, Yingyan (Celine)","Beckmann, Bradford","Ghobadi, Manya"],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-27T19:51:20Z","subjects":[],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/80294","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Krishna, Tushar"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Mahajan, Divya","Lin, Yingyan (Celine)","Beckmann, Bradford","Ghobadi, Manya"]},{"key":"dc:creator","label":"Author","values":["Won, William Jonghoon"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-26T14:58:08Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-26T14:58:08Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-12"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science, PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Georgia Institute of Technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/80294"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Foundation machine learning (ML) models have emerged as one of the most prominent applications in modern computing, exemplified by mixture-of-experts–based large language models. The immense resource demands of these models have driven the development of large-scale, high-performance computing platforms tailored for artificial intelligence workloads. In such distributed platforms, both model parameters and data are partitioned and processed across numerous neural processing units, requiring frequent synchronization of activations and gradients through collective communication operations. As collective communication constitutes a primary bottleneck in distributed ML, optimizing its efficiency remains a critical research challenge. This dissertation explores software-hardware optimizations for collective communications to better understand the tightly coupled design space of networking in distributed ML platforms. First, it introduces ASTRA-sim2.0, an end-to-end simulation and modeling framework that enables comprehensive design space exploration of distributed ML platforms with arbitrary parallelization strategies and multi-dimensional networks. Second, it presents LIBRA, which enhances the bandwidth utilization of hierarchical collective communication algorithms by optimizing multi-dimensional network topologies via analytical modeling. Finally, the dissertation proposes two collective communication algorithm synthesizers, TACOS and PCCL, which automatically generate optimized collective communication algorithms for arbitrary network topologies through algorithmic approaches. Together, the dissertation underscores the significance of judicious software-hardware approaches in achieving efficient collective communication for large-scale distributed ML platforms."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Software-Hardware Optimizations for Efficient Collective Communications in Distributed Machine Learning Platforms"]}]}],"canonical_facts":{"dc:contributor.advisor":["Krishna, Tushar"],"dc:contributor.committeemember":["Mahajan, Divya","Lin, Yingyan (Celine)","Beckmann, Bradford","Ghobadi, Manya"],"dc:creator":["Won, William Jonghoon"],"dc:date.accessioned":["2026-01-26T14:58:08Z"],"dc:date.available":["2026-01-26T14:58:08Z"],"dc:date.issued":["2025-12"],"dc:description.abstract":["Foundation machine learning (ML) models have emerged as one of the most prominent applications in modern computing, exemplified by mixture-of-experts–based large language models. The immense resource demands of these models have driven the development of large-scale, high-performance computing platforms tailored for artificial intelligence workloads. In such distributed platforms, both model parameters and data are partitioned and processed across numerous neural processing units, requiring frequent synchronization of activations and gradients through collective communication operations. As collective communication constitutes a primary bottleneck in distributed ML, optimizing its efficiency remains a critical research challenge. This dissertation explores software-hardware optimizations for collective communications to better understand the tightly coupled design space of networking in distributed ML platforms. First, it introduces ASTRA-sim2.0, an end-to-end simulation and modeling framework that enables comprehensive design space exploration of distributed ML platforms with arbitrary parallelization strategies and multi-dimensional networks. Second, it presents LIBRA, which enhances the bandwidth utilization of hierarchical collective communication algorithms by optimizing multi-dimensional network topologies via analytical modeling. Finally, the dissertation proposes two collective communication algorithm synthesizers, TACOS and PCCL, which automatically generate optimized collective communication algorithms for arbitrary network topologies through algorithmic approaches. Together, the dissertation underscores the significance of judicious software-hardware approaches in achieving efficient collective communication for large-scale distributed ML platforms."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/80294"],"dc:language.iso":["English"],"dc:title":["Software-Hardware Optimizations for Efficient Collective Communications in Distributed Machine Learning Platforms"],"dc:type":["Text"],"thesis:degree_name":["Computer Science, PhD"],"thesis:institution_name":["Georgia Institute of Technology"]},"updated_at":"2026-07-27T19:51:20Z"}