Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 16 of 16 for “"rollback recovery"”.
-
Compiler-assisted rollback recovery
… applying compile-time information to assist in rollback recovery. Depending on the configuration of the original system, one or more of the following three compiler capabilities are applied: (1) code insertion, (2) code analysis, and (3) code transformation.
-
Memory management and rollback recovery in parallel architectures
This thesis examines memory management and rollback recovery in parallel architectures. Three memory management strategies for rapid rollback recovery are studied in this thesis. The first is a cache-based error recovery scheme for shared memory multiprocessors. The second is a design for …
-
Compiler-Assisted Multiple Instruction Rollback Recovery Using a Read Buffer
Multiple instruction rollback (MIR) is a technique to provide rapid recovery from transient processor failures and has been implemented in hardware by researchers and also in mainframe computers. Hardware-based MIR designs eliminate rollback data hazards by providing data redundancy implemented in …
-
Performance evaluation of checkpoint rollback recovery algorithms in distributed systems
Performance evaluation of checkpoint rollback recovery strategies for distributed systems is a field which has not been studied much. Considerable work has been completed in the performance analysis of checkpoint strategies in centralized systems. The necessity for such a study is clear considering …
-
Time-Based Coordinated Checkpointing
A large number of checkpoint-based recovery protocols have been proposed in the literature, however, most of them were never evaluated. The thesis describes the design and implementation of a run-time system for clusters of workstations that allows the rapid testing of checkpoint protocols with …
-
Space reclamation for uncoordinated checkpointing in message-passing systems
Checkpointing and rollback recovery are techniques that can provide efficient recovery from transient process failures. In a message-passing system, the rollback of a message sender may cause the rollback of the corresponding receiver, and the system needs to roll back to a consistent set of …
-
Scalable message-logging techniques for effective fault tolerance in HPC applications
… to incorporate some fashion of fault tolerance. Rollback-recovery techniques provide a framework to overcome crashes in the system by periodically saving the state of the application and rolling back to checkpoints in case of failures. Two well-known rollback-recovery techniques are …
-
An Environment for Imprecise Computations
… deadlines and of integrating checkpointing and rollback recovery with flexible applications as a means to increase the availability and dependability of real-time systems. This thesis makes the following contributions: (1) It describes a model for representing flexible real-time applications as …
-
Failure avoidance techniques for HPC systems based on failure prediction
… popular and used techniques from this field are rollback-recovery protocols. However, existing rollback-recovery techniques have severe scalability limitations and without further optimizations the use of current protocols is put under serious questions for future exascale systems. A way of …
-
Scalable fault tolerance for high-performance streaming dataflow
… systems either come with long downtimes during recovery, or fail to scale to large deployments due to the overhead of global coordination. For example, in the failure of a single dataflow node, existing lineage-based techniques take a long time to recompute all lost and downstream state, while …
-
Toward high-level synthesis of reliable circuits through low-cost modulo shadow datapaths
… result check) for soft errors, we also explore a rollback recovery method with an additional area cost of 28.0% for both cases, observing 411x increase in reliability against soft errors.
-
A fault-tolerant building block for transputer networks for real-time processing
… due to the cost of either synchronization or rollback recovery. The use of redundant live execution images to repair the faulty module guarantees forward fault recoveries.
-
Low-cost error detection through high-level synthesis
… result check) for soft errors, we also explore a rollback recovery method with an additional area cost of 28.0%, observing a 175x increase in reliability against soft errors. Another solution for rapid post-silicon validation of accelerator designs is Hybrid Quick Error Detection (H-QED): …
-
Metamori: A library for Incremental File Checkpointing
The advent of cluster computing has resulted in a thrust towards providing software mechanisms for reliability on clusters. The prevalent model for such mechanisms is to take a snapshot of the state of an application, called a checkpoint and commit it to stable storage. This checkpoint has …
-
Detecting and Recovering from In-Core Hardware Faults Through Software Anomaly Treatment
… relies on hardware support for checkpointing and rollback recovery. When dealing with fault recovery in the presence of I/O, we identify that existing software-level mechanisms that handle output buffering fall short. This dissertation therefore pro- poses a simple low-cost hardware buffer for …
-
Robust and reliable hardware accelerator design through high-level synthesis
… result check) for soft errors, we also explore a rollback recovery method with an additional area cost of 28.0%, observing a 175× increase in reliability against soft errors. While the area cost of our modulo shadow datapaths is much better than traditional modular redundancy approaches, we want …