Abstract
dc:description.abstract<p>As the scale of supercomputers grows, it is becoming increasingly important for software to efficiently withstand hardware and software faults. Process replication is one resilience technique, but typical implementations require replicas to stay closely synchronized with each other. We propose algorithms to lazily detect faults in replicated MPI applications, allowing for more flexibility in replica scheduling and potential power savings. Evaluation shows that, when all processes are operated at full power, this approach allows applications to complete substantially faster as compared to using a synchronized model, and often as fast as in non-replicated execution.</p>
Degree
thesis:*- Name thesis:degree_name
- MS in Computer Science
- Discipline thesis:degree_discipline
- Computer Science
- Year dc:date.available
- 2016
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ford, Corey
- Contributors dc:contributor
-
- Chris Lupo
Identifiers
dc:identifier.*- Identifier
- 10.15368/theses.2016.41
- OAI identifier oai:identifier
- oai:digitalcommons.calpoly.edu:theses-2730