{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/102953"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/102953","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed","abstract":"Supercomputers have played an essential role in the progress of science and engineering research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional application-level checkpointing may no longer be a viable fault tolerance mechanism. This poses serious ramifications for a system's ability to guarantee reliability and availability of its resources. It is becoming increasingly important to understand fault-to-failure propagation and to identify key areas of instrumentation in HPC systems for avoidance, detection, diagnosis, mitigation, and recovery of faults. This thesis presents a software-implemented, prototype-based fault injection tool called HPCArrow and a fault injection methodology as a means to investigate and evaluate HPC application and system resiliency. We demonstrate HPCArrow's capabilities through four fault injection campaigns on a Cray XE/XK hybrid testbed, covering single injections, time-varying or delayed injections, and injections during recovery. These injections emulate failures on network and compute components. The results of these campaigns provide insight into application-level and system-level resiliencies. Across various HPC application frameworks, there are notable deficiencies in fault tolerance. Our experiments also revealed a failure phenomenon that was previously unobserved in field data: application hangs, in which forward progress is not made, but jobs are not terminated until the maximum allowed time has elapsed. At the system level, failover procedures prove highly robust on small-scale systems, able to handle both single and multiple faults in the network.","abstract_html":"Supercomputers have played an essential role in the progress of science and engineering research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional application-level checkpointing may no longer be a viable fault tolerance mechanism. This poses serious ramifications for a system&#x27;s ability to guarantee reliability and availability of its resources. It is becoming increasingly important to understand fault-to-failure propagation and to identify key areas of instrumentation in HPC systems for avoidance, detection, diagnosis, mitigation, and recovery of faults. This thesis presents a software-implemented, prototype-based fault injection tool called HPCArrow and a fault injection methodology as a means to investigate and evaluate HPC application and system resiliency. We demonstrate HPCArrow&#x27;s capabilities through four fault injection campaigns on a Cray XE/XK hybrid testbed, covering single injections, time-varying or delayed injections, and injections during recovery. These injections emulate failures on network and compute components. The results of these campaigns provide insight into application-level and system-level resiliencies. Across various HPC application frameworks, there are notable deficiencies in fault tolerance. Our experiments also revealed a failure phenomenon that was previously unobserved in field data: application hangs, in which forward progress is not made, but jobs are not terminated until the maximum allowed time has elapsed. At the system level, failover procedures prove highly robust on small-scale systems, able to handle both single and multiple faults in the network.","abstract_has_math":false,"creators":["Tang, Sharon S."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Kalbarczyk, Zbigniew T."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-02-08T18:44:42Z","date_published":"2019-02-08T18:44:42Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Resiliency","Reliability","Fault Injections","High-Performance Computing","Interconnects"],"languages":["en"],"rights":["Copyright 2018 Sharon S. Tang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/102953","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kalbarczyk, Zbigniew T."]},{"key":"dc:creator","label":"Author","values":["Tang, Sharon S."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-02-08T18:44:42Z","2021-02-09T10:15:30Z","2018-12-11","2018-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Resiliency","Reliability","Fault Injections","High-Performance Computing","Interconnects"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Sharon S. Tang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/102953"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Supercomputers have played an essential role in the progress of science and engineering research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional application-level checkpointing may no longer be a viable fault tolerance mechanism. This poses serious ramifications for a system's ability to guarantee reliability and availability of its resources. It is becoming increasingly important to understand fault-to-failure propagation and to identify key areas of instrumentation in HPC systems for avoidance, detection, diagnosis, mitigation, and recovery of faults. This thesis presents a software-implemented, prototype-based fault injection tool called HPCArrow and a fault injection methodology as a means to investigate and evaluate HPC application and system resiliency. We demonstrate HPCArrow's capabilities through four fault injection campaigns on a Cray XE/XK hybrid testbed, covering single injections, time-varying or delayed injections, and injections during recovery. These injections emulate failures on network and compute components. The results of these campaigns provide insight into application-level and system-level resiliencies. Across various HPC application frameworks, there are notable deficiencies in fault tolerance. Our experiments also revealed a failure phenomenon that was previously unobserved in field data: application hangs, in which forward progress is not made, but jobs are not terminated until the maximum allowed time has elapsed. At the system level, failover procedures prove highly robust on small-scale systems, able to handle both single and multiple faults in the network.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-12-01","The student, Sharon Tang, accepted the attached license on 2018-12-11 at 10:07.","The student, Sharon Tang, submitted this Thesis for approval on 2018-12-11 at 10:08.","This Thesis was approved for publication on 2018-12-11 at 10:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13272 on 2019-02-08 at 11:41:43","Made available in DSpace on 2019-02-08T18:44:42Z (GMT). No. of bitstreams: 2 TANG-THESIS-2018.pdf: 11311816 bytes, checksum: 5a53c04b42b0457a4607fcfa4b17e556 (MD5) LICENSE.txt: 4208 bytes, checksum: ee32e4a3e6de6b47c87e6380052158de (MD5) Previous issue date: 2018-12-11","Embargo set by: Seth Robbins for item 109981 Lift date: 2021-02-08T18:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 109981 on 2021-02-09T10:15:30Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed"]}]}],"canonical_facts":{"dc:contributor":["Kalbarczyk, Zbigniew T."],"dc:creator":["Tang, Sharon S."],"dc:date":["2019-02-08T18:44:42Z","2021-02-09T10:15:30Z","2018-12-11","2018-12"],"dc:description":["Supercomputers have played an essential role in the progress of science and engineering research. As the high-performance computing (HPC) community moves towards the next generation of HPC computing, it faces several challenges, one of which is reliability of HPC systems. Error rates are expected to significantly increase on exascale systems to the point where traditional application-level checkpointing may no longer be a viable fault tolerance mechanism. This poses serious ramifications for a system's ability to guarantee reliability and availability of its resources. It is becoming increasingly important to understand fault-to-failure propagation and to identify key areas of instrumentation in HPC systems for avoidance, detection, diagnosis, mitigation, and recovery of faults. This thesis presents a software-implemented, prototype-based fault injection tool called HPCArrow and a fault injection methodology as a means to investigate and evaluate HPC application and system resiliency. We demonstrate HPCArrow's capabilities through four fault injection campaigns on a Cray XE/XK hybrid testbed, covering single injections, time-varying or delayed injections, and injections during recovery. These injections emulate failures on network and compute components. The results of these campaigns provide insight into application-level and system-level resiliencies. Across various HPC application frameworks, there are notable deficiencies in fault tolerance. Our experiments also revealed a failure phenomenon that was previously unobserved in field data: application hangs, in which forward progress is not made, but jobs are not terminated until the maximum allowed time has elapsed. At the system level, failover procedures prove highly robust on small-scale systems, able to handle both single and multiple faults in the network.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-12-01","The student, Sharon Tang, accepted the attached license on 2018-12-11 at 10:07.","The student, Sharon Tang, submitted this Thesis for approval on 2018-12-11 at 10:08.","This Thesis was approved for publication on 2018-12-11 at 10:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13272 on 2019-02-08 at 11:41:43","Made available in DSpace on 2019-02-08T18:44:42Z (GMT). No. of bitstreams: 2 TANG-THESIS-2018.pdf: 11311816 bytes, checksum: 5a53c04b42b0457a4607fcfa4b17e556 (MD5) LICENSE.txt: 4208 bytes, checksum: ee32e4a3e6de6b47c87e6380052158de (MD5) Previous issue date: 2018-12-11","Embargo set by: Seth Robbins for item 109981 Lift date: 2021-02-08T18:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 109981 on 2021-02-09T10:15:30Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/102953"],"dc:language":["en"],"dc:rights":["Copyright 2018 Sharon S. Tang"],"dc:subject":["Resiliency","Reliability","Fault Injections","High-Performance Computing","Interconnects"],"dc:title":["Resiliency of high-performance computing systems: A fault-injection-based characterization of the high-speed network in the blue waters testbed"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}