Massachusetts Institute of Technology
A container-based lightweight fault tolerance framework for high performance computing workloads
Abstract
dc:description.abstractAccording to the latest world's top 500 supercomputers list, ~90% of the top High Performance Computing (HPC) systems are based on commodity hardware clusters, which are typically designed for performance rather than reliability. The Mean Time Between Failures (MTBF) for some current petascale systems has been reported to be several days, while studies estimate it may be less than 60 minutes for future exascale systems. One of the largest studies on HPC system failures showed that more than 50% of failures were due to hardware, and that failure rates grew with system size. Hence, running extended workloads on such systems is becoming more challenging as system sizes grow. In this work, we design and implement a lightweight fault tolerance framework to improve the sustainability of running workloads on HPC clusters. The framework mainly includes a fault prediction component and a remedy component.
Degree
thesis:*- Name thesis:degree_name
- Doctoral
- Department dc:contributor.department
- Massachusetts Institute of Technology. Department of Civil and Environmental Engineering
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2019
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sindi, Mohamad(Mohamad Othman)
- Advisor dc:contributor.advisor
-
- John R. Williams.
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1721.1/124188
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/124188