Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 70 for “"checkpointing"”.

  1. Time-Based Coordinated Checkpointing

    A large number of checkpoint-based recovery protocols have been proposed in the literature, however, most of them were never evaluated. The thesis describes the design and implementation of a run-time system for clusters of workstations that allows the rapid testing of checkpoint protocols with …

    uiuc Repository record for Time-Based Coordinated Checkpointing (opens in a new tab)

  2. Cooperative checkpointing for supercomputing systems

    A system-level checkpointing mechanism, with global knowledge of the state and health of the machine, can improve performance and reliability by dynamically deciding when to skip checkpoint requests made by applications. This thesis presents such a technique, called cooperative checkpointing, and …

    mit Repository record for Cooperative checkpointing for supercomputing systems (opens in a new tab)

  3. Scalable Diskless Checkpointing for Large Parallel Systems

    … to investigate the performability of diskless checkpointing. Our model evaluation shows that the overhead of checkpoint/recovery is small on systems with thousands of nodes, and with appropriate partitioning of nodes, the user application can survive several times longer.

    uiuc Repository record for Scalable Diskless Checkpointing for Large Parallel Systems (opens in a new tab)

  4. Rebound: Scalable checkpointing for coherent shared memory

    … to large manycores, the hardware-based global checkpointing schemes that have been proposed for small shared-memory machines do not scale. Scalability barriers include global operations, work lost to global rollback, and inefficiencies in imbalanced or I/O-intensive loads. Scalable …

    uiuc Repository record for Rebound: Scalable checkpointing for coherent shared memory (opens in a new tab)

  5. Metamori: A library for Incremental File Checkpointing

    … be checkpointed. Several libraries exist for checkpointing volatile state. Some of these libraries feature incremental checkpointing, where only the changes since the last checkpoint are recorded in the next checkpoint. Such incremental checkpointing is advantageous since otherwise, the time …

    vt Repository record for Metamori: A library for Incremental File Checkpointing (opens in a new tab)

  6. Checkpointing of parallel applications in a Grid environment

    … reliably with minimum overhead incurred. While checkpointing is the most common method to achieve fault tolerance, there is still a lot of work to be done to improve the efficiency of the mechanism. This thesis provides an in-depth description of a novel solution for checkpointing parallel …

    westminster Repository record for Checkpointing of parallel applications in a Grid environment (opens in a new tab)

  7. Security for proof of work blockchains via checkpointing

    Finality gadgets are comprised of a Byzantine Fault Tolerant (BFT) protocol finalizing blocks produced by a Proof-of-Work (PoW) or Proof-of-Stake (PoS) chain protocol. They have become very popular methods for combining the best features of the BFT and PoW protocols and are proposed for deployment …

    uiuc Repository record for Security for proof of work blockchains via checkpointing (opens in a new tab)

  8. Concurrent checkpointing for fast recovery in object-based systems

    … a single instruction and lower than that of checkpointing.

    uiuc Repository record for Concurrent checkpointing for fast recovery in object-based systems (opens in a new tab)

  9. Space reclamation for uncoordinated checkpointing in message-passing systems

    Checkpointing and rollback recovery are techniques that can provide efficient recovery from transient process failures. In a message-passing system, the rollback of a message sender may cause the rollback of the corresponding receiver, and the system needs to roll back to a consistent set of …

    uiuc Repository record for Space reclamation for uncoordinated checkpointing in message-passing systems (opens in a new tab)

  10. An investigation of semantic checkpointing in dynamically reconfigurable applications

    … dynamic reconfiguration by enabling semantic checkpointing: capturing application level state. The reconfiguration under question is not restricted to replacement of a few modules or moving modules from one machine to another. The individual components can be completely changed as long as the …

    mit Repository record for An investigation of semantic checkpointing in dynamically reconfigurable applications (opens in a new tab)

  11. Extending Scojo-PECT by migration based on application level checkpointing

    … the gained execution time minus overhead of checkpointing/migration, to make comparison with original execution time.

    windsor Repository record for Extending Scojo-PECT by migration based on application level checkpointing (opens in a new tab)

  12. Recuperação com base em Checkpointing : uma abordagem orientada a objetos

    … programador, através do uso de uma biblioteca de checkpointing que implemente os mecanismos de salvamento de estados e recuperação é o ponto focal deste trabalho. Diante do contexto exposto, nesta dissertação, são definidas e implementadas as classes de uma biblioteca que provê mecanismos de …

    brazil-ufrgs Repository record for Recuperação com base em Checkpointing : uma abordagem orientada a objetos (opens in a new tab)

  13. Checkpointing in distributed virtual memory by utilizing local virtual memory

    This study explores a recovery strategy using checkpointing in a distributed shared virtual memory (DVM) system. DVM shares virtual memory in a loosely-coupled multi-computer system and is implemented at the software-level. The goal of this recovery strategy is to obtain a consistent recovery line …

    uiuc Repository record for Checkpointing in distributed virtual memory by utilizing local virtual memory (opens in a new tab)

  14. Fault Tolerance for Real-Time Systems: Analysis and Optimization of Roll-back Recovery with Checkpointing

    … tolerance technique, Roll-back Recovery with Checkpointing (RRC), that efficiently copes with soft errors. The major drawback of RRC is that it introduces a time overhead which depends on the number of checkpoints that are used in RRC. Depending on how the checkpoints are distributed …

    lund Repository record for Fault Tolerance for Real-Time Systems: Analysis and Optimization of Roll-back Recovery with Checkpointing (opens in a new tab)

  15. Heterogeneous and Mobile Recovery

    … control maximum recovery time for log-based checkpointing protocols and may cause unnecessary checkpointing overhead. In this thesis, an adaptive checkpointing protocol is designed to accurately enforce the user-defined recovery time and to reduce excessive checkpoints. Instead of using fixed …

    uiuc Repository record for Heterogeneous and Mobile Recovery (opens in a new tab)

  16. A semantic checkpoint framework for enabling runtime-reconfigurable applications

    … applications through the use of semantic checkpointing. We view applications here as a collection of inter-connected components, and reconfigurations as the reconstitution of components that make up an application. By checkpointing only values that are deemed to be of semantic …

    mit Repository record for A semantic checkpoint framework for enabling runtime-reconfigurable applications (opens in a new tab)

  17. FastRecover: simple and effective fault recovery in a distributed operator-based stream processing engine

    … tolerance mechanism: the overhead of additional checkpointing operations during normal processing, and the time required to recover and return to normal processing when a failure happens. The main challenge lies in that faster recovery requires higher checkpointing overhead, and vice versa. This …

    uiuc Repository record for FastRecover: simple and effective fault recovery in a distributed operator-based stream processing engine (opens in a new tab)

  18. Investigating system resilience in distributed evolutionary GAN training

    … We apply a combination of uncoordinated checkpointing, attempted reconnecting, and restarting nodes to form a simple and efficient solution for system resilience in Lipizzaner. We find that checkpointing and reconnecting are essential and simple solutions to failure recovery in …

    mit Repository record for Investigating system resilience in distributed evolutionary GAN training (opens in a new tab)

  19. Discrete Adjoints: Theoretical Analysis, Efficient Computation, and Applications

    … habilitation presents an analysis of different checkpointing strategies to reduce the memory requirement of the discrete adjoint computation. Additionally, a new algorithm for computing sparse Hessian matrices is presented including a complexity analysis and a report on practical experiments. …

    qucosa-diss

  20. Resource Allocation in Vehicular Cloud Computing

    … and also in the case where cars may leave while checkpointing is in progress, leading to system failure. A comprehensive set of simulations have shown that our theoretical predictions are accurate.</p> <p>We considered two different environments for scheduling strategy: deterministic and …

    odu Repository record for Resource Allocation in Vehicular Cloud Computing (opens in a new tab)

Page 1 of 4