{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/386642"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/386642","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Separating conflict-recovery from failure-recovery in distributed consensus","abstract":"Distributed databases provide fault tolerance while allowing multiple clients to concurrently submit requests and have those requests executed as if they are performed using a single thread on a single machine. This makes them vital components for a variety of distributed systems, from banking transactions to managing cluster configurations. At the core of these distributed databases is their consensus protocol which must order concurrently submitted requests and ensure that the system can mask or recover from failures. Current state-of-the-art consensus protocols re-use their failure-recovery mechanism to resolve ordering conflicts. The thesis of this dissertation is that it is feasible to separate the mechanisms for conflict-recovery and failure-recovery, and that doing so is effective in allowing each to be optimised separately. To support this thesis we first categorise existing conflict-recovery mechanisms and introduce Multi-Shot-FastPaxos which provides a uniform interface to express, and in some cases further optimise, conflict-recovery mechanisms separately from failure-recovery mechanisms. We also propose a new failure masking approach that reduces the overhead of masking coordinator failures and evaluate existing mechanisms for recovery from coordinator failure, allowing us to propose an optimised recovery procedure for etcd. Finally we present a case study designing a new protocol, Unanimous 2 Phase Commit, which from an initial basic implementation is optimised to equal or exceed the steady-state performance of the current state-of-the-art on all workloads, with an optimised recovery procedure.","abstract_html":"Distributed databases provide fault tolerance while allowing multiple clients to concurrently submit requests and have those requests executed as if they are performed using a single thread on a single machine. This makes them vital components for a variety of distributed systems, from banking transactions to managing cluster configurations. At the core of these distributed databases is their consensus protocol which must order concurrently submitted requests and ensure that the system can mask or recover from failures. Current state-of-the-art consensus protocols re-use their failure-recovery mechanism to resolve ordering conflicts. The thesis of this dissertation is that it is feasible to separate the mechanisms for conflict-recovery and failure-recovery, and that doing so is effective in allowing each to be optimised separately. To support this thesis we first categorise existing conflict-recovery mechanisms and introduce Multi-Shot-FastPaxos which provides a uniform interface to express, and in some cases further optimise, conflict-recovery mechanisms separately from failure-recovery mechanisms. We also propose a new failure masking approach that reduces the overhead of masking coordinator failures and evaluate existing mechanisms for recovery from coordinator failure, allowing us to propose an optimised recovery procedure for etcd. Finally we present a case study designing a new protocol, Unanimous 2 Phase Commit, which from an initial basic implementation is optimised to equal or exceed the steady-state performance of the current state-of-the-art on all workloads, with an optimised recovery procedure.","abstract_has_math":false,"creators":["Jensen, Christopher"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Mortier, Richard","Howard, Heidi"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-01","date_published":"2025-01","updated_at":"2026-07-24T01:33:25Z","subjects":["Consensus","Distributed Systems","Paxos","FastPaxos","Failure","Benchmark"],"languages":[],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/1349cc66-553f-42d5-af1a-21a7a6c3e190/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.119774","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Mortier, Richard","Howard, Heidi"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Huawei"]},{"key":"dc:creator","label":"Author","values":["Jensen, Christopher"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-01"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/386642"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Consensus","Distributed Systems","Paxos","FastPaxos","Failure","Benchmark"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/1349cc66-553f-42d5-af1a-21a7a6c3e190/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.119774"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/a3418429-13ba-4c0e-accb-d00e57ed946b/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Distributed databases provide fault tolerance while allowing multiple clients to concurrently submit requests and have those requests executed as if they are performed using a single thread on a single machine. This makes them vital components for a variety of distributed systems, from banking transactions to managing cluster configurations. At the core of these distributed databases is their consensus protocol which must order concurrently submitted requests and ensure that the system can mask or recover from failures. Current state-of-the-art consensus protocols re-use their failure-recovery mechanism to resolve ordering conflicts. The thesis of this dissertation is that it is feasible to separate the mechanisms for conflict-recovery and failure-recovery, and that doing so is effective in allowing each to be optimised separately. To support this thesis we first categorise existing conflict-recovery mechanisms and introduce Multi-Shot-FastPaxos which provides a uniform interface to express, and in some cases further optimise, conflict-recovery mechanisms separately from failure-recovery mechanisms. We also propose a new failure masking approach that reduces the overhead of masking coordinator failures and evaluate existing mechanisms for recovery from coordinator failure, allowing us to propose an optimised recovery procedure for etcd. Finally we present a case study designing a new protocol, Unanimous 2 Phase Commit, which from an initial basic implementation is optimised to equal or exceed the steady-state performance of the current state-of-the-art on all workloads, with an optimised recovery procedure."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["0ea35b36a7bbcf1b1a0be770b424d7e5","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Separating conflict-recovery from failure-recovery in distributed consensus"]}]}],"canonical_facts":{"dc:contributor.advisor":["Mortier, Richard","Howard, Heidi"],"dc:contributor.sponsor":["Huawei"],"dc:creator":["Jensen, Christopher"],"dc:date.issued":["2025-01"],"dc:description.abstract":["Distributed databases provide fault tolerance while allowing multiple clients to concurrently submit requests and have those requests executed as if they are performed using a single thread on a single machine. This makes them vital components for a variety of distributed systems, from banking transactions to managing cluster configurations. At the core of these distributed databases is their consensus protocol which must order concurrently submitted requests and ensure that the system can mask or recover from failures. Current state-of-the-art consensus protocols re-use their failure-recovery mechanism to resolve ordering conflicts. The thesis of this dissertation is that it is feasible to separate the mechanisms for conflict-recovery and failure-recovery, and that doing so is effective in allowing each to be optimised separately. To support this thesis we first categorise existing conflict-recovery mechanisms and introduce Multi-Shot-FastPaxos which provides a uniform interface to express, and in some cases further optimise, conflict-recovery mechanisms separately from failure-recovery mechanisms. We also propose a new failure masking approach that reduces the overhead of masking coordinator failures and evaluate existing mechanisms for recovery from coordinator failure, allowing us to propose an optimised recovery procedure for etcd. Finally we present a case study designing a new protocol, Unanimous 2 Phase Commit, which from an initial basic implementation is optimised to equal or exceed the steady-state performance of the current state-of-the-art on all workloads, with an optimised recovery procedure."],"dc:format.checksum.md5":["0ea35b36a7bbcf1b1a0be770b424d7e5","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.119774"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/a3418429-13ba-4c0e-accb-d00e57ed946b/download"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/386642"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/1349cc66-553f-42d5-af1a-21a7a6c3e190/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"dc:subject":["Consensus","Distributed Systems","Paxos","FastPaxos","Failure","Benchmark"],"dc:title":["Separating conflict-recovery from failure-recovery in distributed consensus"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T01:33:25Z"}