{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101213"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101213","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Topology-aware distributed graph processing for tightly-coupled clusters","abstract":"Cloud applications have burgeoned over the last few years, but they are typically written for loosely-coupled clusters such as datacenters. In this thesis we investigate how one can run cloud applications in tightly-coupled clusters and network topologies, namely super-computers. Specifically, we look at a class of distributed machine learning systems called distributed graph processing systems, and run them on NCSA Blue Waters. Partitioning the graph is key to achieving performance in distributed graph processing systems. We present new topology-aware partitioning techniques that better exploit the structure of the network topologies in supercomputers. Compared to existing work, our new Restricted Oblivious and Grid Centroid partitioning approaches produce 25-33% improvement in makespan, along with a sizable reduction in network traffic. We also discuss optimizations such as smart network buffers that further amplify the improvement. To help operators select the best graph partitioning technique, we culminate our experimental results into a decision tree.","abstract_html":"Cloud applications have burgeoned over the last few years, but they are typically written for loosely-coupled clusters such as datacenters. In this thesis we investigate how one can run cloud applications in tightly-coupled clusters and network topologies, namely super-computers. Specifically, we look at a class of distributed machine learning systems called distributed graph processing systems, and run them on NCSA Blue Waters. Partitioning the graph is key to achieving performance in distributed graph processing systems. We present new topology-aware partitioning techniques that better exploit the structure of the network topologies in supercomputers. Compared to existing work, our new Restricted Oblivious and Grid Centroid partitioning approaches produce 25-33% improvement in makespan, along with a sizable reduction in network traffic. We also discuss optimizations such as smart network buffers that further amplify the improvement. To help operators select the best graph partitioning technique, we culminate our experimental results into a decision tree.","abstract_has_math":false,"creators":["Bhatt, Mayank"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gupta, Indranil"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:36:52Z","date_published":"2018-09-04T20:36:52Z","updated_at":"2026-07-22T22:24:38Z","subjects":["graph processing","supercomputers","topology"],"languages":["en"],"rights":["Copyright 2018 Mayank Bhatt"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101213","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gupta, Indranil"]},{"key":"dc:creator","label":"Author","values":["Bhatt, Mayank"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:36:52Z","2020-09-05T09:15:32Z","2018-04-24","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["graph processing","supercomputers","topology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Mayank Bhatt"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101213"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Cloud applications have burgeoned over the last few years, but they are typically written for loosely-coupled clusters such as datacenters. In this thesis we investigate how one can run cloud applications in tightly-coupled clusters and network topologies, namely super-computers. Specifically, we look at a class of distributed machine learning systems called distributed graph processing systems, and run them on NCSA Blue Waters. Partitioning the graph is key to achieving performance in distributed graph processing systems. We present new topology-aware partitioning techniques that better exploit the structure of the network topologies in supercomputers. Compared to existing work, our new Restricted Oblivious and Grid Centroid partitioning approaches produce 25-33% improvement in makespan, along with a sizable reduction in network traffic. We also discuss optimizations such as smart network buffers that further amplify the improvement. To help operators select the best graph partitioning technique, we culminate our experimental results into a decision tree.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Mayank Bhatt, accepted the attached license on 2018-04-23 at 17:13.","The student, Mayank Bhatt, submitted this Thesis for approval on 2018-04-23 at 17:20.","This Thesis was approved for publication on 2018-04-24 at 15:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12435 on 2018-08-31 at 17:21:19","Made available in DSpace on 2018-09-04T20:36:52Z (GMT). No. of bitstreams: 2 BHATT-THESIS-2018.pdf: 1415794 bytes, checksum: e08311d8168967b2e47baf1ef67f7fdc (MD5) LICENSE.txt: 4209 bytes, checksum: b810a770b0873fc45062dd7e9ce83fde (MD5) Previous issue date: 2018-04-24","Embargo set by: Seth Robbins for item 107297 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107297 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107297 on 2020-09-05T09:15:32Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Topology-aware distributed graph processing for tightly-coupled clusters"]}]}],"canonical_facts":{"dc:contributor":["Gupta, Indranil"],"dc:creator":["Bhatt, Mayank"],"dc:date":["2018-09-04T20:36:52Z","2020-09-05T09:15:32Z","2018-04-24","2018-05"],"dc:description":["Cloud applications have burgeoned over the last few years, but they are typically written for loosely-coupled clusters such as datacenters. In this thesis we investigate how one can run cloud applications in tightly-coupled clusters and network topologies, namely super-computers. Specifically, we look at a class of distributed machine learning systems called distributed graph processing systems, and run them on NCSA Blue Waters. Partitioning the graph is key to achieving performance in distributed graph processing systems. We present new topology-aware partitioning techniques that better exploit the structure of the network topologies in supercomputers. Compared to existing work, our new Restricted Oblivious and Grid Centroid partitioning approaches produce 25-33% improvement in makespan, along with a sizable reduction in network traffic. We also discuss optimizations such as smart network buffers that further amplify the improvement. To help operators select the best graph partitioning technique, we culminate our experimental results into a decision tree.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Mayank Bhatt, accepted the attached license on 2018-04-23 at 17:13.","The student, Mayank Bhatt, submitted this Thesis for approval on 2018-04-23 at 17:20.","This Thesis was approved for publication on 2018-04-24 at 15:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12435 on 2018-08-31 at 17:21:19","Made available in DSpace on 2018-09-04T20:36:52Z (GMT). No. of bitstreams: 2 BHATT-THESIS-2018.pdf: 1415794 bytes, checksum: e08311d8168967b2e47baf1ef67f7fdc (MD5) LICENSE.txt: 4209 bytes, checksum: b810a770b0873fc45062dd7e9ce83fde (MD5) Previous issue date: 2018-04-24","Embargo set by: Seth Robbins for item 107297 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107297 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107297 on 2020-09-05T09:15:32Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101213"],"dc:language":["en"],"dc:rights":["Copyright 2018 Mayank Bhatt"],"dc:subject":["graph processing","supercomputers","topology"],"dc:title":["Topology-aware distributed graph processing for tightly-coupled clusters"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}