Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 70 for “"checkpointing"”.
-
Time-Based Coordinated Checkpointing
A large number of checkpoint-based recovery protocols have been proposed in the literature, however, most of them were never evaluated. The thesis describes the design and implementation of a run-time system for clusters of workstations that allows the rapid testing of checkpoint protocols with …
-
Cooperative checkpointing for supercomputing systems
A system-level checkpointing mechanism, with global knowledge of the state and health of the machine, can improve performance and reliability by dynamically deciding when to skip checkpoint requests made by applications. This thesis presents such a technique, called cooperative checkpointing, and …
-
Scalable Diskless Checkpointing for Large Parallel Systems
… to investigate the performability of diskless checkpointing. Our model evaluation shows that the overhead of checkpoint/recovery is small on systems with thousands of nodes, and with appropriate partitioning of nodes, the user application can survive several times longer.
-
Rebound: Scalable checkpointing for coherent shared memory
… to large manycores, the hardware-based global checkpointing schemes that have been proposed for small shared-memory machines do not scale. Scalability barriers include global operations, work lost to global rollback, and inefficiencies in imbalanced or I/O-intensive loads. Scalable …
-
Metamori: A library for Incremental File Checkpointing
… be checkpointed. Several libraries exist for checkpointing volatile state. Some of these libraries feature incremental checkpointing, where only the changes since the last checkpoint are recorded in the next checkpoint. Such incremental checkpointing is advantageous since otherwise, the time …
-
Checkpointing of parallel applications in a Grid environment
… reliably with minimum overhead incurred. While checkpointing is the most common method to achieve fault tolerance, there is still a lot of work to be done to improve the efficiency of the mechanism. This thesis provides an in-depth description of a novel solution for checkpointing parallel …
-
Security for proof of work blockchains via checkpointing
Finality gadgets are comprised of a Byzantine Fault Tolerant (BFT) protocol finalizing blocks produced by a Proof-of-Work (PoW) or Proof-of-Stake (PoS) chain protocol. They have become very popular methods for combining the best features of the BFT and PoW protocols and are proposed for deployment …
-
Concurrent checkpointing for fast recovery in object-based systems
… a single instruction and lower than that of checkpointing.
-
Space reclamation for uncoordinated checkpointing in message-passing systems
Checkpointing and rollback recovery are techniques that can provide efficient recovery from transient process failures. In a message-passing system, the rollback of a message sender may cause the rollback of the corresponding receiver, and the system needs to roll back to a consistent set of …
-
An investigation of semantic checkpointing in dynamically reconfigurable applications
… dynamic reconfiguration by enabling semantic checkpointing: capturing application level state. The reconfiguration under question is not restricted to replacement of a few modules or moving modules from one machine to another. The individual components can be completely changed as long as the …
-
Extending Scojo-PECT by migration based on application level checkpointing
… the gained execution time minus overhead of checkpointing/migration, to make comparison with original execution time.
-
Recuperação com base em Checkpointing : uma abordagem orientada a objetos
… programador, através do uso de uma biblioteca de checkpointing que implemente os mecanismos de salvamento de estados e recuperação é o ponto focal deste trabalho. Diante do contexto exposto, nesta dissertação, são definidas e implementadas as classes de uma biblioteca que provê mecanismos de …
-
Checkpointing in distributed virtual memory by utilizing local virtual memory
This study explores a recovery strategy using checkpointing in a distributed shared virtual memory (DVM) system. DVM shares virtual memory in a loosely-coupled multi-computer system and is implemented at the software-level. The goal of this recovery strategy is to obtain a consistent recovery line …
-
Fault Tolerance for Real-Time Systems: Analysis and Optimization of Roll-back Recovery with Checkpointing
… tolerance technique, Roll-back Recovery with Checkpointing (RRC), that efficiently copes with soft errors. The major drawback of RRC is that it introduces a time overhead which depends on the number of checkpoints that are used in RRC. Depending on how the checkpoints are distributed …
-
Heterogeneous and Mobile Recovery
… control maximum recovery time for log-based checkpointing protocols and may cause unnecessary checkpointing overhead. In this thesis, an adaptive checkpointing protocol is designed to accurately enforce the user-defined recovery time and to reduce excessive checkpoints. Instead of using fixed …
-
A semantic checkpoint framework for enabling runtime-reconfigurable applications
… applications through the use of semantic checkpointing. We view applications here as a collection of inter-connected components, and reconfigurations as the reconstitution of components that make up an application. By checkpointing only values that are deemed to be of semantic …
-
FastRecover: simple and effective fault recovery in a distributed operator-based stream processing engine
… tolerance mechanism: the overhead of additional checkpointing operations during normal processing, and the time required to recover and return to normal processing when a failure happens. The main challenge lies in that faster recovery requires higher checkpointing overhead, and vice versa. This …
-
Investigating system resilience in distributed evolutionary GAN training
… We apply a combination of uncoordinated checkpointing, attempted reconnecting, and restarting nodes to form a simple and efficient solution for system resilience in Lipizzaner. We find that checkpointing and reconnecting are essential and simple solutions to failure recovery in …
-
Discrete Adjoints: Theoretical Analysis, Efficient Computation, and Applications
… habilitation presents an analysis of different checkpointing strategies to reduce the memory requirement of the discrete adjoint computation. Additionally, a new algorithm for computing sparse Hessian matrices is presented including a complexity analysis and a report on practical experiments. …
-
Resource Allocation in Vehicular Cloud Computing
… and also in the case where cars may leave while checkpointing is in progress, leading to system failure. A comprehensive set of simulations have shown that our theoretical predictions are accurate.</p> <p>We considered two different environments for scheduling strategy: deterministic and …
Page 1 of 4