{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/72761"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/72761","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Software topological message aggregation techniques for large-scale parallel systems","abstract":"High overhead of fine-grained communication is a significant performance bottleneck for many classes of applications written for large scale parallel systems. This thesis explores techniques for reducing this overhead through topological aggregation, in which fine-grained messages are dynamically combined not just at each source of communication, but also at intermediate points between source and destination processors. The performance characteristics of aggregation and selection of virtual topology for routing are analyzed theoretically and experimentally. Schemes that exploit fast intra-node communication to reduce the number of inter-node messages are also explored. The presented techniques are implemented for the Charm++ parallel programming system in the Topological Routing and Aggregation Module (TRAM). It is also demonstrated how TRAM can be automated at the level of the runtime system to function with little or no configuration or input required from the user. Using TRAM, the performance of a number of benchmark and scientific applications is evaluated, with speedups of up to 10x for communication benchmarks and up to 4x for full-fledged scientific applications on petascale systems.","abstract_html":"High overhead of fine-grained communication is a significant performance bottleneck for many classes of applications written for large scale parallel systems. This thesis explores techniques for reducing this overhead through topological aggregation, in which fine-grained messages are dynamically combined not just at each source of communication, but also at intermediate points between source and destination processors. The performance characteristics of aggregation and selection of virtual topology for routing are analyzed theoretically and experimentally. Schemes that exploit fast intra-node communication to reduce the number of inter-node messages are also explored. The presented techniques are implemented for the Charm++ parallel programming system in the Topological Routing and Aggregation Module (TRAM). It is also demonstrated how TRAM can be automated at the level of the runtime system to function with little or no configuration or input required from the user. Using TRAM, the performance of a number of benchmark and scientific applications is evaluated, with speedups of up to 10x for communication benchmarks and up to 4x for full-fledged scientific applications on petascale systems.","abstract_has_math":false,"creators":["Wesolowski, Lukasz"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Kale, Laxmikant V.","Gropp, William D.","Torrellas, Josep","Balaji, Pavan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-01-21T19:47:49Z","date_published":"2015-01-21T19:47:49Z","updated_at":"2026-07-22T22:26:07Z","subjects":["Communication Optimization","Message Aggregation"],"languages":["en"],"rights":["Copyright 2014 Lukasz Wesolowski"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/72761","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kale, Laxmikant V.","Gropp, William D.","Torrellas, Josep","Balaji, Pavan"]},{"key":"dc:creator","label":"Author","values":["Wesolowski, Lukasz"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-01-21T19:47:49Z","2014-12","2015-01-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Communication Optimization","Message Aggregation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2014 Lukasz Wesolowski"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/72761"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["High overhead of fine-grained communication is a significant performance bottleneck for many classes of applications written for large scale parallel systems. This thesis explores techniques for reducing this overhead through topological aggregation, in which fine-grained messages are dynamically combined not just at each source of communication, but also at intermediate points between source and destination processors. The performance characteristics of aggregation and selection of virtual topology for routing are analyzed theoretically and experimentally. Schemes that exploit fast intra-node communication to reduce the number of inter-node messages are also explored. The presented techniques are implemented for the Charm++ parallel programming system in the Topological Routing and Aggregation Module (TRAM). It is also demonstrated how TRAM can be automated at the level of the runtime system to function with little or no configuration or input required from the user. Using TRAM, the performance of a number of benchmark and scientific applications is evaluated, with speedups of up to 10x for communication benchmarks and up to 4x for full-fledged scientific applications on petascale systems.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2014-12-02T13:53:05Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Wesolowski_Lukasz.pdf: 2605405 bytes, checksum: 472193eef01a1dc87a4287bd8e76a665 (MD5)","Made available in DSpace on 2015-01-21T19:47:49Z (GMT). No. of bitstreams: 1 Lukasz_Wesolowski.pdf: 2605407 bytes, checksum: d1e93929e284253c73fa61c4cb05028d (MD5)"]},{"key":"dc:title","label":"Title","values":["Software topological message aggregation techniques for large-scale parallel systems"]}]}],"canonical_facts":{"dc:contributor":["Kale, Laxmikant V.","Gropp, William D.","Torrellas, Josep","Balaji, Pavan"],"dc:creator":["Wesolowski, Lukasz"],"dc:date":["2015-01-21T19:47:49Z","2014-12","2015-01-21"],"dc:description":["High overhead of fine-grained communication is a significant performance bottleneck for many classes of applications written for large scale parallel systems. This thesis explores techniques for reducing this overhead through topological aggregation, in which fine-grained messages are dynamically combined not just at each source of communication, but also at intermediate points between source and destination processors. The performance characteristics of aggregation and selection of virtual topology for routing are analyzed theoretically and experimentally. Schemes that exploit fast intra-node communication to reduce the number of inter-node messages are also explored. The presented techniques are implemented for the Charm++ parallel programming system in the Topological Routing and Aggregation Module (TRAM). It is also demonstrated how TRAM can be automated at the level of the runtime system to function with little or no configuration or input required from the user. Using TRAM, the performance of a number of benchmark and scientific applications is evaluated, with speedups of up to 10x for communication benchmarks and up to 4x for full-fledged scientific applications on petascale systems.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2014-12-02T13:53:05Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Wesolowski_Lukasz.pdf: 2605405 bytes, checksum: 472193eef01a1dc87a4287bd8e76a665 (MD5)","Made available in DSpace on 2015-01-21T19:47:49Z (GMT). No. of bitstreams: 1 Lukasz_Wesolowski.pdf: 2605407 bytes, checksum: d1e93929e284253c73fa61c4cb05028d (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/72761"],"dc:language":["en"],"dc:rights":["Copyright 2014 Lukasz Wesolowski"],"dc:subject":["Communication Optimization","Message Aggregation"],"dc:title":["Software topological message aggregation techniques for large-scale parallel systems"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:07Z"}