{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/24127"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/24127","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Detecting and Recovering from In-Core Hardware Faults Through Software Anomaly Treatment","abstract":"Aggressive scaling of CMOS transistors has enabled extensive system integration and building faster and more eﬃcient systems. On the ﬂip side, this has resulted in an increasing number of devices that fail in shipped components in-the-ﬁeld for a variety of reasons including soft errors, wear-out failures, and infant mortality. The pervasiveness of the problem across a broad market demands low cost and generic reliability solutions, precluding traditional solutions that employed excessive redundancy or piecemeal solutions that address only a few failure modes. This dissertation presents SWAT (SoftWare Anomaly Treatment), a low cost resiliency solution that eﬀectively handles hardware faults while incurring low cost during the common mode of fault-free operations. SWAT is based on two key observations about the design of resilient systems. First, only those hardware faults that aﬀect software need to be handled and second, since the common mode of operation is fault-free, fault-free execution should incur near-zero overheads. SWAT thus uses novel zero to low cost hardware and software monitors that watch for anomalous software behavior to detect hardware faults. SWAT then relies on hardware support for checkpointing and rollback recovery. When dealing with fault recovery in the presence of I/O, we identify that existing software-level mechanisms that handle output buﬀering fall short. This dissertation therefore pro- poses a simple low-cost hardware buﬀer for output buﬀering and demonstrates that this strategy achieves high recoverability while incurring low overheads. Although not detailed in this dissertation, SWAT contains a comprehensive diagnosis procedure that is invoked in the rare event of a fault to isolate the root-cause of the fault by distinguishing between software bugs, transient hardware faults, and permanent hardware faults. Eﬀectively, SWAT handles hardware faults uniformly as software bugs, amortizing the resiliency cost across both hardware and software reliability. The results in this dissertation show that the SWAT strategy is eﬀective to detect and recover the system from a variety of in-core permanent and transient faults in various microarchitecture units for both compute-intensive and I/O-intensive workloads. In particular, this dissertation demonstrates that the SWAT detectors detect nearly all permanent and transient faults in most hardware units in both types of workloads, with only a small fraction of the faults corrupting application output.(Certain hardware structures like the FPU may need additional support to be amenable to software anomaly detection.) Further, a majority of these faults are tolerated by the applications due to their inherent fault-tolerant nature, resulting in only 0.2% of the injected faults aﬀecting the application and yielding incorrect outputs (such faults are classiﬁed as Silent Data Corruptions, or SDCs). When attempting to recover the detected faults, we show that handling I/O is important for fault recovery. With our proposed low-cost hardware for output buﬀering, we show that over 94% of the detected faults are recoverable with low performance and area overheads during fault-free execution even in the presence of I/O. Finally, this dissertation builds a fundamental understanding behind why the SWAT strategy is eﬀective for handling faults in modern workloads. The key insight is that the SWAT detectors are adept at detecting perturbations in control operations and memory addresses and a majority of the application values aﬀect such operations. Faults in values that that never aﬀect such operations are hard-to-detect and require additional support to be amenable to software anomaly detection. In summary, this dissertation presents SWAT as a complete solution to detect and recover from from in-core hardware faults. The techniques presented here therefore have far reaching implications on the design of low-cost solutions to handle unreliable hardware.","abstract_html":"Aggressive scaling of CMOS transistors has enabled extensive system integration and building faster and more eﬃcient systems. On the ﬂip side, this has resulted in an increasing number of devices that fail in shipped components in-the-ﬁeld for a variety of reasons including soft errors, wear-out failures, and infant mortality. The pervasiveness of the problem across a broad market demands low cost and generic reliability solutions, precluding traditional solutions that employed excessive redundancy or piecemeal solutions that address only a few failure modes. This dissertation presents SWAT (SoftWare Anomaly Treatment), a low cost resiliency solution that eﬀectively handles hardware faults while incurring low cost during the common mode of fault-free operations. SWAT is based on two key observations about the design of resilient systems. First, only those hardware faults that aﬀect software need to be handled and second, since the common mode of operation is fault-free, fault-free execution should incur near-zero overheads. SWAT thus uses novel zero to low cost hardware and software monitors that watch for anomalous software behavior to detect hardware faults. SWAT then relies on hardware support for checkpointing and rollback recovery. When dealing with fault recovery in the presence of I/O, we identify that existing software-level mechanisms that handle output buﬀering fall short. This dissertation therefore pro- poses a simple low-cost hardware buﬀer for output buﬀering and demonstrates that this strategy achieves high recoverability while incurring low overheads. Although not detailed in this dissertation, SWAT contains a comprehensive diagnosis procedure that is invoked in the rare event of a fault to isolate the root-cause of the fault by distinguishing between software bugs, transient hardware faults, and permanent hardware faults. Eﬀectively, SWAT handles hardware faults uniformly as software bugs, amortizing the resiliency cost across both hardware and software reliability. The results in this dissertation show that the SWAT strategy is eﬀective to detect and recover the system from a variety of in-core permanent and transient faults in various microarchitecture units for both compute-intensive and I/O-intensive workloads. In particular, this dissertation demonstrates that the SWAT detectors detect nearly all permanent and transient faults in most hardware units in both types of workloads, with only a small fraction of the faults corrupting application output.(Certain hardware structures like the FPU may need additional support to be amenable to software anomaly detection.) Further, a majority of these faults are tolerated by the applications due to their inherent fault-tolerant nature, resulting in only 0.2% of the injected faults aﬀecting the application and yielding incorrect outputs (such faults are classiﬁed as Silent Data Corruptions, or SDCs). When attempting to recover the detected faults, we show that handling I/O is important for fault recovery. With our proposed low-cost hardware for output buﬀering, we show that over 94% of the detected faults are recoverable with low performance and area overheads during fault-free execution even in the presence of I/O. Finally, this dissertation builds a fundamental understanding behind why the SWAT strategy is eﬀective for handling faults in modern workloads. The key insight is that the SWAT detectors are adept at detecting perturbations in control operations and memory addresses and a majority of the application values aﬀect such operations. Faults in values that that never aﬀect such operations are hard-to-detect and require additional support to be amenable to software anomaly detection. In summary, this dissertation presents SWAT as a complete solution to detect and recover from from in-core hardware faults. The techniques presented here therefore have far reaching implications on the design of low-cost solutions to handle unreliable hardware.","abstract_has_math":false,"creators":["Ramachandran, Pradeep"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Adve, Sarita V.","Adve, Vikram S.","King, Samuel T.","Snir, Marc","Bose, Pradip"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-25T14:51:14Z","date_published":"2011-05-25T14:51:14Z","updated_at":"2026-07-22T22:25:23Z","subjects":["Fault tolerance","Computer architecture","symptom detection"],"languages":["en"],"rights":["Copyright 2011 Pradeep Ramachandran"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/24127","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Adve, Sarita V.","Adve, Vikram S.","King, Samuel T.","Snir, Marc","Bose, Pradip"]},{"key":"dc:creator","label":"Author","values":["Ramachandran, Pradeep"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-25T14:51:14Z","2011-05"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Fault tolerance","Computer architecture","symptom detection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Pradeep Ramachandran"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/24127"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Aggressive scaling of CMOS transistors has enabled extensive system integration and building faster and more eﬃcient systems. On the ﬂip side, this has resulted in an increasing number of devices that fail in shipped components in-the-ﬁeld for a variety of reasons including soft errors, wear-out failures, and infant mortality. The pervasiveness of the problem across a broad market demands low cost and generic reliability solutions, precluding traditional solutions that employed excessive redundancy or piecemeal solutions that address only a few failure modes. This dissertation presents SWAT (SoftWare Anomaly Treatment), a low cost resiliency solution that eﬀectively handles hardware faults while incurring low cost during the common mode of fault-free operations. SWAT is based on two key observations about the design of resilient systems. First, only those hardware faults that aﬀect software need to be handled and second, since the common mode of operation is fault-free, fault-free execution should incur near-zero overheads. SWAT thus uses novel zero to low cost hardware and software monitors that watch for anomalous software behavior to detect hardware faults. SWAT then relies on hardware support for checkpointing and rollback recovery. When dealing with fault recovery in the presence of I/O, we identify that existing software-level mechanisms that handle output buﬀering fall short. This dissertation therefore pro- poses a simple low-cost hardware buﬀer for output buﬀering and demonstrates that this strategy achieves high recoverability while incurring low overheads. Although not detailed in this dissertation, SWAT contains a comprehensive diagnosis procedure that is invoked in the rare event of a fault to isolate the root-cause of the fault by distinguishing between software bugs, transient hardware faults, and permanent hardware faults. Eﬀectively, SWAT handles hardware faults uniformly as software bugs, amortizing the resiliency cost across both hardware and software reliability. The results in this dissertation show that the SWAT strategy is eﬀective to detect and recover the system from a variety of in-core permanent and transient faults in various microarchitecture units for both compute-intensive and I/O-intensive workloads. In particular, this dissertation demonstrates that the SWAT detectors detect nearly all permanent and transient faults in most hardware units in both types of workloads, with only a small fraction of the faults corrupting application output.(Certain hardware structures like the FPU may need additional support to be amenable to software anomaly detection.) Further, a majority of these faults are tolerated by the applications due to their inherent fault-tolerant nature, resulting in only 0.2% of the injected faults aﬀecting the application and yielding incorrect outputs (such faults are classiﬁed as Silent Data Corruptions, or SDCs). When attempting to recover the detected faults, we show that handling I/O is important for fault recovery. With our proposed low-cost hardware for output buﬀering, we show that over 94% of the detected faults are recoverable with low performance and area overheads during fault-free execution even in the presence of I/O. Finally, this dissertation builds a fundamental understanding behind why the SWAT strategy is eﬀective for handling faults in modern workloads. The key insight is that the SWAT detectors are adept at detecting perturbations in control operations and memory addresses and a majority of the application values aﬀect such operations. Faults in values that that never aﬀect such operations are hard-to-detect and require additional support to be amenable to software anomaly detection. In summary, this dissertation presents SWAT as a complete solution to detect and recover from from in-core hardware faults. The techniques presented here therefore have far reaching implications on the design of low-cost solutions to handle unreliable hardware.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-04-18T20:49:52Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Ramachandran_Pradeep.pdf: 1939067 bytes, checksum: 6da5a540fe3f34306ef652d7bd4d7516 (MD5)","Made available in DSpace on 2011-05-25T14:51:14Z (GMT). No. of bitstreams: 2 Ramachandran_Pradeep.pdf: 1939067 bytes, checksum: 6da5a540fe3f34306ef652d7bd4d7516 (MD5) license.txt: 4070 bytes, checksum: 1e563a5c302325c6c85812528c4c20de (MD5)"]},{"key":"dc:title","label":"Title","values":["Detecting and Recovering from In-Core Hardware Faults Through Software Anomaly Treatment"]}]}],"canonical_facts":{"dc:contributor":["Adve, Sarita V.","Adve, Vikram S.","King, Samuel T.","Snir, Marc","Bose, Pradip"],"dc:creator":["Ramachandran, Pradeep"],"dc:date":["2011-05-25T14:51:14Z","2011-05"],"dc:description":["Aggressive scaling of CMOS transistors has enabled extensive system integration and building faster and more eﬃcient systems. On the ﬂip side, this has resulted in an increasing number of devices that fail in shipped components in-the-ﬁeld for a variety of reasons including soft errors, wear-out failures, and infant mortality. The pervasiveness of the problem across a broad market demands low cost and generic reliability solutions, precluding traditional solutions that employed excessive redundancy or piecemeal solutions that address only a few failure modes. This dissertation presents SWAT (SoftWare Anomaly Treatment), a low cost resiliency solution that eﬀectively handles hardware faults while incurring low cost during the common mode of fault-free operations. SWAT is based on two key observations about the design of resilient systems. First, only those hardware faults that aﬀect software need to be handled and second, since the common mode of operation is fault-free, fault-free execution should incur near-zero overheads. SWAT thus uses novel zero to low cost hardware and software monitors that watch for anomalous software behavior to detect hardware faults. SWAT then relies on hardware support for checkpointing and rollback recovery. When dealing with fault recovery in the presence of I/O, we identify that existing software-level mechanisms that handle output buﬀering fall short. This dissertation therefore pro- poses a simple low-cost hardware buﬀer for output buﬀering and demonstrates that this strategy achieves high recoverability while incurring low overheads. Although not detailed in this dissertation, SWAT contains a comprehensive diagnosis procedure that is invoked in the rare event of a fault to isolate the root-cause of the fault by distinguishing between software bugs, transient hardware faults, and permanent hardware faults. Eﬀectively, SWAT handles hardware faults uniformly as software bugs, amortizing the resiliency cost across both hardware and software reliability. The results in this dissertation show that the SWAT strategy is eﬀective to detect and recover the system from a variety of in-core permanent and transient faults in various microarchitecture units for both compute-intensive and I/O-intensive workloads. In particular, this dissertation demonstrates that the SWAT detectors detect nearly all permanent and transient faults in most hardware units in both types of workloads, with only a small fraction of the faults corrupting application output.(Certain hardware structures like the FPU may need additional support to be amenable to software anomaly detection.) Further, a majority of these faults are tolerated by the applications due to their inherent fault-tolerant nature, resulting in only 0.2% of the injected faults aﬀecting the application and yielding incorrect outputs (such faults are classiﬁed as Silent Data Corruptions, or SDCs). When attempting to recover the detected faults, we show that handling I/O is important for fault recovery. With our proposed low-cost hardware for output buﬀering, we show that over 94% of the detected faults are recoverable with low performance and area overheads during fault-free execution even in the presence of I/O. Finally, this dissertation builds a fundamental understanding behind why the SWAT strategy is eﬀective for handling faults in modern workloads. The key insight is that the SWAT detectors are adept at detecting perturbations in control operations and memory addresses and a majority of the application values aﬀect such operations. Faults in values that that never aﬀect such operations are hard-to-detect and require additional support to be amenable to software anomaly detection. In summary, this dissertation presents SWAT as a complete solution to detect and recover from from in-core hardware faults. The techniques presented here therefore have far reaching implications on the design of low-cost solutions to handle unreliable hardware.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-04-18T20:49:52Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Ramachandran_Pradeep.pdf: 1939067 bytes, checksum: 6da5a540fe3f34306ef652d7bd4d7516 (MD5)","Made available in DSpace on 2011-05-25T14:51:14Z (GMT). No. of bitstreams: 2 Ramachandran_Pradeep.pdf: 1939067 bytes, checksum: 6da5a540fe3f34306ef652d7bd4d7516 (MD5) license.txt: 4070 bytes, checksum: 1e563a5c302325c6c85812528c4c20de (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/24127"],"dc:language":["en"],"dc:rights":["Copyright 2011 Pradeep Ramachandran"],"dc:subject":["Fault tolerance","Computer architecture","symptom detection"],"dc:title":["Detecting and Recovering from In-Core Hardware Faults Through Software Anomaly Treatment"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:23Z"}