Universität Passau
Improving Performance of Byzantine State Machine Replication through Self-Adaptation
Abstract
dc:description.abstractByzantine Fault-Tolerant (BFT) State Machine Replication (SMR) provides strong resilience against faults and intrusions, making it a solid foundation for building dependable distributed systems. However, its practical adoption remains limited due to several challenges: high operational costs stemming from the additional replicas required for fault tolerance, performance overhead from increased synchronization and communication, and the complexity of configuring and tuning consensus protocols. This thesis aims to address these limitations by introducing self-adaptation capabilities into BFT SMR systems to improve performance and reduce configuration complexity. The research is guided by three central questions. First, how BFT SMR systems can be made self-adaptive, enabling them to dynamically optimize their configuration and behavior in response to changing environments and workloads. Second, how such dynamic optimizations can be efficiently and systematically evaluated to assess their effectiveness and discover optimal configurations for specific scenarios. Finally, what performance potentials can be realized through self-optimization, particularly in terms of reducing latency in wide-area deployments. The first contribution realizes BFT data streaming by integrating the data streaming platform Kafka with the BFT SMR middleware BFT-SMaRt, resulting in a practically usable and resilient data streaming system. It shows how BFT SMR can exploit the specific properties of data streaming - namely, append-only state semantics and a partitionable application state - to enable optimizations such as speculative request execution and parallel processing across state partitions. These techniques help preserve the scalability of data streaming while reducing the latency overhead introduced by the multi-step agreement protocol, especially in high-latency environments. The second contribution is an extensible optimization framework that integrates self-adaptation capabilities to a wide range of current and future BFT SMR systems. It introduces an architecture centered around a self-adaptive feedback control loop, enabling continuous monitoring and automated reconfiguration. The framework supports the definition of scenario-driven optimization strategies by mapping high-level goals to relevant monitoring and control parameters. It also includes protocols that deterministically govern the optimization process while ensuring resilient and non-intrusive system adaptation. The third contribution is an automated evaluation platform that enables large-scale testing of distributed systems within a virtualized environment. It automates the provisioning of computational resources across multiple cloud platforms, and supports the emulation of realistic conditions, including wide-area network scenarios. The platform allows the definition of repeatable workloads and fault scenarios through declarative playbooks. Resulting experiments are fully self-contained and can be executed in any compatible environment, ensuring that evaluations are portable, reproducible, and verifiable. The fourth contribution of this thesis is a network federation approach that enables inter-cloud operations through the establishment of private overlay networks. Based on an extensive evaluation of VPN technologies, it introduces an agent-based federation mechanism that leverages WireGuard and GRE to provide flexible and transparent Layer2 or Layer3 overlay networks, while preserving data integrity and confidentiality. Prototype evaluations confirm its efficiency and portability across diverse cloud platforms. The final contribution of this thesis is a location-awareness concept for BFT SMR systems deployed in wide-area networks, where response times are adversely affected by the geographical distance between clients and server nodes. To address this, the proposed latency optimization makes replicas aware of their geographic location and allows this location to be configured at runtime. Through self-adaptation and continuous latency monitoring, replica groups can assess the impact of their placement on overall request latency and autonomously initiate relocations toward regions with higher client activity. An evaluation of the prototype implementation demonstrates a significant reduction of the average request latency experienced by clients.
Degree
thesis:*- Level thesis:degree_level
- thesis.doctoral
- Grantor dc:publisher
- Universität Passau
- Year
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Köstler, Johannes
- Contributors dc:contributor
-
- Reiser, Hans P.
- Becker, Christian
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Standardbedingung laut Einverständniserklärung
Identifiers
dc:identifier.*- Repository record source_url
- https://opus4.kobv.de/opus4-uni-passau/frontdoor/index/index/docId/2069
- OAI identifier oai:identifier
- oai:kobv.de-opus4-uni-passau:2069