Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 14 of 14 for “"COLLECTIVE COMMUNICATION"”.

  1. Unresponsiveness -Tolerant Collective Communication

    … the final goal of increasing the performance of collective-communication-centric parallel applications.

    uiuc Repository record for Unresponsiveness -Tolerant Collective Communication (opens in a new tab)

  2. The Collective Communication of Social Choice Messages

    … in this dissertation is to develop a theory of collective communication. Collective communication is defined as social interaction mediated through messages whose production involves a collectivity. The focus of analysis is on social choice messages, messages that prescribe or proscribe the …

    penn Repository record for The Collective Communication of Social Choice Messages (opens in a new tab)

  3. HARP: A MACHINE LEARNING FRAMEWORK ON TOP OF THE COLLECTIVE COMMUNICATION LAYER FOR THE BIG DATA SOFTWARE STACK

    … In this thesis, the Harp framework, which uses collective communication techniques, is prototyped to improve the performance of data movement and provides high-level APIs for various synchronization patterns in iterative computation. In contrast to traditional parallelization strategies that …

    iu Repository record for HARP: A MACHINE LEARNING FRAMEWORK ON TOP OF THE COLLECTIVE COMMUNICATION LAYER FOR THE BIG DATA SOFTWARE STACK (opens in a new tab)

  4. Software-Hardware Optimizations for Efficient Collective Communications in Distributed Machine Learning Platforms

    … of activations and gradients through collective communication operations. As collective communication constitutes a primary bottleneck in distributed ML, optimizing its efficiency remains a critical research challenge. This dissertation explores software-hardware optimizations for …

    gatech Repository record for Software-Hardware Optimizations for Efficient Collective Communications in Distributed Machine Learning Platforms (opens in a new tab)

  5. Accelerating MPI collective communications through hierarchical algorithms with flexible inter-node communication and imbalance awareness

    … work presents and evaluates algorithms for MPI collective communication operations on high performance systems. Collective communication algorithms are extensively investigated, and a universal algorithm to improve the performance of MPI collective operations on hierarchical clusters is …

    purdue-thes Repository record for Accelerating MPI collective communications through hierarchical algorithms with flexible inter-node communication and imbalance awareness (opens in a new tab)

  6. On Communication-Computation Overlap in High-Performance Computing

    … to be kept at an absolute minimum, including communication operations. Non-blocking-collective operations extend the concept of collective operations by offering the additional benefit of being able to overlap communication and computation. However, it has been demonstrated that collective

    houston Repository record for On Communication-Computation Overlap in High-Performance Computing (opens in a new tab)

  7. Resource Sharing for Machine Learning Serving

    … such as computation kernels, memory usage, and collective communication, are struggling to keep pace with the increasingly integrated, irregular, and massive machine learning models. This dissertation proposes resource sharing as a fundamental design principle to address these emerging …

    penn Repository record for Resource Sharing for Machine Learning Serving (opens in a new tab)

  8. Improving the performance of parallel scientific applications using cache injection

    … cache injection policy, and the application's communication characteristics. Cache injection addresses the memory wall for I/O by writing data into a processor's cache directly from the I/O bus. This technique, unlike data prefetching, reduces the number of reads served by the memory unit. This …

    unm Repository record for Improving the performance of parallel scientific applications using cache injection (opens in a new tab)

  9. Performance Tuning and Modeling of Communication in Parallel Applications

    … view, namely, the computational section and the communication section. The time spent in communication operations is a major factor in determining the scalability of parallel applications. Tuning the parameters of a communication library can be used to adapt its characteristics to a particular …

    houston Repository record for Performance Tuning and Modeling of Communication in Parallel Applications (opens in a new tab)

  10. OS INSPIRED COMPLETE KERNEL FUSION

    … Learning (DML) is increasingly recognized as a communication-bound workload, with most existing work aiming to alleviate this bottleneck through communication-computation overlap. However, current approaches—largely reliant on CPU-managed operator scheduling and synchronous communication

    cornell Repository record for OS INSPIRED COMPLETE KERNEL FUSION (opens in a new tab)

  11. Configurable Hierarchical Allreduce Algorithms for HPC Systems

    Collective communication is a critical component of large scale scientific applications across all message size ranges. However, optimizing performance in the small to medium message regime is especially challenging because both bandwidth and latency influence efficiency. Many workloads, including …

    uic

  12. High-Performance Network- and GPU-Aware Communication for MPI Partitioned and MPI Neighbourhoods

    … are poorly optimized for multi-threaded communication, and MPI currently has no standardized way to enable GPU support. To address some of the issues with MPI, MPI Partitioned Point-to-Point Communication was proposed to better support hybrid workloads. We addressed the lack of …

    queens Repository record for High-Performance Network- and GPU-Aware Communication for MPI Partitioned and MPI Neighbourhoods (opens in a new tab)

  13. Simulation-based performance analysis and tuning for future supercomputers

    Hardware and software co-design is becoming increasingly important due to complexities in supercomputing architectures. Simulating applications before there is access to the real hardware can assist machine architects in making better design decisions that can optimize application performance. At …

    uiuc Repository record for Simulation-based performance analysis and tuning for future supercomputers (opens in a new tab)

  14. La contribution du leadership à la construction de l'intelligence collective dans la production d'un bulletin de nouvelles télévisé

    … leadership à la construction de l’intelligence collective dans une équipe de travail qui produit un bulletin de nouvelles télévisé. L’intelligence collective est une façon de travailler qu’ont développée les organisations hautement fiables, c’est-à-dire les organisations où la moindre erreur …

    ottawa-retro Repository record for La contribution du leadership à la construction de l'intelligence collective dans la production d'un bulletin de nouvelles télévisé (opens in a new tab)