Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 52 for “"HPC systems"”.

  1. Declarative Analytics on Heterogeneous HPC Systems

    The emergence of exascale systems marks a transforming era in high-performance computing (HPC) powered by extensive use of GPUs. GPGPU's popularity in HPC, due to performance gains and power efficiency, demands redesigning traditional algorithms to exploit GPU parallelism. However, declarative …

    uic

  2. Configurable Hierarchical Allreduce Algorithms for HPC Systems

    … hierarchy present in current high performance systems, where intra node communication differs significantly from inter node communication. This thesis introduces a configurable hierarchical design for the Allgather and Reduce Scatter collectives, which serve as the building blocks for a …

    uic

  3. Failure avoidance techniques for HPC systems based on failure prediction

    … in today's large high-performance computing systems is wasted due to failures and recoveries. Moreover, it is expected that high performance computing will reach exascale within a decade, decreasing the mean time between failures to one day or even a few hours, making fault tolerance a major …

    uiuc Repository record for Failure avoidance techniques for HPC systems based on failure prediction (opens in a new tab)

  4. Common-case optimized memory hierarchy for data centers and HPC systems

    … data centers and high performance computing (HPC) systems; as such, it is time to rethink memory system designs. Conventional memory system designs in existing systems often seek to provide uniform performance across time and space. While this design approach is simple, which benefits hardware …

    uiuc Repository record for Common-case optimized memory hierarchy for data centers and HPC systems (opens in a new tab)

  5. Mitigating variability in HPC systems and applications for performance and power efficiency

    … large-scale High Performance Computing (HPC) data centers. For example, current production petaflop supercomputers consume more than 10 megawatts of machine and cooling power that costs millions of dollars every year. As HPC moves towards exascale computing, these costs will increase and …

    uiuc Repository record for Mitigating variability in HPC systems and applications for performance and power efficiency (opens in a new tab)

  6. Modular high-speed signaling for networks in HPC systems using COTS components

    Thesis (S.B. and M. Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2000.

    mit Repository record for Modular high-speed signaling for networks in HPC systems using COTS components (opens in a new tab)

  7. Towards Using Free Memory to Improve Microarchitecture Performance

    … Through a large-scale study of four production HPC systems, we find that memory underutilization problem in HPC systems is very severe. As unused memory is wasted memory, we propose exposing a compute node's unused memory to its CPU(s) through a user-transparent CPU-OS codesign. This can enable …

    vt Repository record for Towards Using Free Memory to Improve Microarchitecture Performance (opens in a new tab)

  8. Towards a Resource Efficient Framework for Distributed Deep Learning Applications

    … (GPUs) in latest high-performance computing (HPC) supercomputing systems. HPC architectures bring different performance trends in training throughput compared to the existing studies. Multiple GPUs and high-speed interconnect are used for distributed deep learning on HPC systems. Extant …

    vt Repository record for Towards a Resource Efficient Framework for Distributed Deep Learning Applications (opens in a new tab)

  9. Automated Runtime Analysis and Adaptation for Scalable Heterogeneous Computing

    … consequence, current high-performance computing (HPC) systems integrate a wide variety of compute resources with different capabilities and execution models, ranging from multi-core CPUs to many-core accelerators. While such heterogeneous systems can enable dramatic acceleration of user …

    vt Repository record for Automated Runtime Analysis and Adaptation for Scalable Heterogeneous Computing (opens in a new tab)

  10. Using Rollback Avoidance to Mitigate Failures in Next-Generation Extreme-Scale Systems

    High-performance computing (HPC) systems enable scientists to numerically model complex phenomena in many important physical systems. The next major milestone in the development of HPC systems is the construction of the first supercomputer capable executing more than an exaflop, 10^18 floating …

    unm Repository record for Using Rollback Avoidance to Mitigate Failures in Next-Generation Extreme-Scale Systems (opens in a new tab)

  11. Auto-tuned optimized parallel I/O for GIScience and spatial applications

    … bottleneck of using high-performance computing systems as we are heading towards the Exascale era. An unprecedented amount of data is being produced everyday by different sources. On the other hand, the computation power of HPC systems is getting scaled to hundreds of thousands cores. However, …

    uiuc Repository record for Auto-tuned optimized parallel I/O for GIScience and spatial applications (opens in a new tab)

  12. Models and Techniques for Green High-Performance Computing

    High-performance computing (HPC) systems have become power limited. For instance, the U.S. Department of Energy set a power envelope of 20MW in 2008 for the first exascale supercomputer now expected to arrive in 2021--22. Toward this end, we seek to improve the greenness of HPC systems by improving …

    vt Repository record for Models and Techniques for Green High-Performance Computing (opens in a new tab)

  13. Prediction Models for Multi-dimensional Power-Performance Optimization on Many Cores

    Power has become a primary concern for HPC systems. Dynamic voltage and frequency scaling (DVFS) and dynamic concurrency throttling (DCT) are two software tools (or knobs) for reducing the dynamic power consumption of HPC systems. To date, few works have considered the synergistic integration of …

    vt Repository record for Prediction Models for Multi-dimensional Power-Performance Optimization on Many Cores (opens in a new tab)

  14. Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed

    … research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional …

    uiuc Repository record for Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed (opens in a new tab)

  15. Energy Measurements of High Performance Computing Systems: From Instrumentation to Analysis

    … energy measurement solution for production HPC systems and address its shortcomings. Such high-resolution and large-scale measurements present challenges regarding the management of large volumes of generated metric data. I address these challenges with a scalable infrastructure for …

    qucosa-diss

  16. An Adaptive Framework for Managing Heterogeneous Many-Core Clusters

    … such applications, High-Performance Computing (HPC) systems need to employ thousands of cores and innovative data management. At the same time, an emerging trend in designing HPC systems is to leverage specialized asymmetric multicores, such as IBM Cell and AMD Fusion APUs, and commodity …

    vt Repository record for An Adaptive Framework for Managing Heterogeneous Many-Core Clusters (opens in a new tab)

  17. A container-based lightweight fault tolerance framework for high performance computing workloads

    … ~90% of the top High Performance Computing (HPC) systems are based on commodity hardware clusters, which are typically designed for performance rather than reliability. The Mean Time Between Failures (MTBF) for some current petascale systems has been reported to be several days, while studies …

    mit Repository record for A container-based lightweight fault tolerance framework for high performance computing workloads (opens in a new tab)

  18. Using High-Performance Computing to Scale Generative Adversarial Networks

    … tested at scale in high performance computing(HPC) systems. We believe that by utilizing HPC technologies, we can scale up Lipizzaner and observe performance enhancements. This thesis achieves this scale up, using Oak Ridge National Labs’ Summit Supercomputer. We observed improvements in the …

    mit Repository record for Using High-Performance Computing to Scale Generative Adversarial Networks (opens in a new tab)

  19. Evaluating the Impact of GPU Frequency Tuning and Power Capping on Performance and Efficiency

    SUMMARY High-Performance Computing (HPC) systems play a pivotal role in modern scientific and en- gineering advancements. However, the increasing demand for computational power comes with significant energy consumption challenges. This thesis investigates the impact of GPU fre- quency tuning and …

    uic

Page 1 of 3