Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 52 for “"HPC systems"”.
-
Declarative Analytics on Heterogeneous HPC Systems
The emergence of exascale systems marks a transforming era in high-performance computing (HPC) powered by extensive use of GPUs. GPGPU's popularity in HPC, due to performance gains and power efficiency, demands redesigning traditional algorithms to exploit GPU parallelism. However, declarative …
-
Configurable Hierarchical Allreduce Algorithms for HPC Systems
… hierarchy present in current high performance systems, where intra node communication differs significantly from inter node communication. This thesis introduces a configurable hierarchical design for the Allgather and Reduce Scatter collectives, which serve as the building blocks for a …
-
Failure avoidance techniques for HPC systems based on failure prediction
… in today's large high-performance computing systems is wasted due to failures and recoveries. Moreover, it is expected that high performance computing will reach exascale within a decade, decreasing the mean time between failures to one day or even a few hours, making fault tolerance a major …
-
Common-case optimized memory hierarchy for data centers and HPC systems
… data centers and high performance computing (HPC) systems; as such, it is time to rethink memory system designs. Conventional memory system designs in existing systems often seek to provide uniform performance across time and space. While this design approach is simple, which benefits hardware …
-
Mitigating variability in HPC systems and applications for performance and power efficiency
… large-scale High Performance Computing (HPC) data centers. For example, current production petaflop supercomputers consume more than 10 megawatts of machine and cooling power that costs millions of dollars every year. As HPC moves towards exascale computing, these costs will increase and …
-
Modular high-speed signaling for networks in HPC systems using COTS components
Thesis (S.B. and M. Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2000.
-
Towards Using Free Memory to Improve Microarchitecture Performance
… Through a large-scale study of four production HPC systems, we find that memory underutilization problem in HPC systems is very severe. As unused memory is wasted memory, we propose exposing a compute node's unused memory to its CPU(s) through a user-transparent CPU-OS codesign. This can enable …
-
Towards a Resource Efficient Framework for Distributed Deep Learning Applications
… (GPUs) in latest high-performance computing (HPC) supercomputing systems. HPC architectures bring different performance trends in training throughput compared to the existing studies. Multiple GPUs and high-speed interconnect are used for distributed deep learning on HPC systems. Extant …
-
Automated Runtime Analysis and Adaptation for Scalable Heterogeneous Computing
… consequence, current high-performance computing (HPC) systems integrate a wide variety of compute resources with different capabilities and execution models, ranging from multi-core CPUs to many-core accelerators. While such heterogeneous systems can enable dramatic acceleration of user …
-
Using Rollback Avoidance to Mitigate Failures in Next-Generation Extreme-Scale Systems
High-performance computing (HPC) systems enable scientists to numerically model complex phenomena in many important physical systems. The next major milestone in the development of HPC systems is the construction of the first supercomputer capable executing more than an exaflop, 10^18 floating …
-
Auto-tuned optimized parallel I/O for GIScience and spatial applications
… bottleneck of using high-performance computing systems as we are heading towards the Exascale era. An unprecedented amount of data is being produced everyday by different sources. On the other hand, the computation power of HPC systems is getting scaled to hundreds of thousands cores. However, …
-
Models and Techniques for Green High-Performance Computing
High-performance computing (HPC) systems have become power limited. For instance, the U.S. Department of Energy set a power envelope of 20MW in 2008 for the first exascale supercomputer now expected to arrive in 2021--22. Toward this end, we seek to improve the greenness of HPC systems by improving …
-
Prediction Models for Multi-dimensional Power-Performance Optimization on Many Cores
Power has become a primary concern for HPC systems. Dynamic voltage and frequency scaling (DVFS) and dynamic concurrency throttling (DCT) are two software tools (or knobs) for reducing the dynamic power consumption of HPC systems. To date, few works have considered the synergistic integration of …
-
Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed
… research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional …
-
Energy Measurements of High Performance Computing Systems: From Instrumentation to Analysis
… energy measurement solution for production HPC systems and address its shortcomings. Such high-resolution and large-scale measurements present challenges regarding the management of large volumes of generated metric data. I address these challenges with a scalable infrastructure for …
-
A Framework for Efficient Management of Fault Tolerance in Cloud Data Centres and High-Performance Computing Systems: An Investigation and Performance analysis of a Cloud Based Virtual Machine Success and Failure Rate in a typical Cloud Computing Environment and Prediction Methods
… to be an issue of growing concern in cloud and HPC systems, mitigating the impact of failure and providing accurate predictions with enough lead time remains a difficult research problem. Traditional existing fault-tolerance strategies such as regular check-point/restart and replication are not …
-
An Adaptive Framework for Managing Heterogeneous Many-Core Clusters
… such applications, High-Performance Computing (HPC) systems need to employ thousands of cores and innovative data management. At the same time, an emerging trend in designing HPC systems is to leverage specialized asymmetric multicores, such as IBM Cell and AMD Fusion APUs, and commodity …
-
A container-based lightweight fault tolerance framework for high performance computing workloads
… ~90% of the top High Performance Computing (HPC) systems are based on commodity hardware clusters, which are typically designed for performance rather than reliability. The Mean Time Between Failures (MTBF) for some current petascale systems has been reported to be several days, while studies …
-
Using High-Performance Computing to Scale Generative Adversarial Networks
… tested at scale in high performance computing(HPC) systems. We believe that by utilizing HPC technologies, we can scale up Lipizzaner and observe performance enhancements. This thesis achieves this scale up, using Oak Ridge National Labs’ Summit Supercomputer. We observed improvements in the …
-
Evaluating the Impact of GPU Frequency Tuning and Power Capping on Performance and Efficiency
SUMMARY High-Performance Computing (HPC) systems play a pivotal role in modern scientific and en- gineering advancements. However, the increasing demand for computational power comes with significant energy consumption challenges. This thesis investigates the impact of GPU fre- quency tuning and …
Page 1 of 3