University of Exeter
Trustworthy Federated Learning Systems: From Secure Distributed Training to Reliable Fine-tuning
Abstract
dc:descriptionTraining and fine-tuning Deep Learning (DL) models require vast amounts of domain-specific data. However, this data is often distributed across different organizations and devices, restricted from direct sharing by privacy, ownership, and regulatory constraints. While Federated Learning (FL) emerged to facilitate collaborative training without exposing raw data, it introduces new trust and security vulnerabilities. Specifically, malicious clients can extract sensitive information from model updates or inject poisoned updates to degrade the global model's performance. Furthermore, the widespread adoption of Large Language Models (LLMs) has expanded the threat landscape well beyond the training phase. Because LLMs process unpredictable user inputs and unverified external content, they are susceptible to novel inference-time vulnerabilities. These include jailbreaking and prompt injection, which can fundamentally alter model behavior. To address these problems, this thesis makes three contributions. First, it proposes a blockchain-secured Federated Learning framework with a Zero-Knowledge Proof of Training (ZKPoT) consensus mechanism that enables verifiable and privacy-preserving coordination among distributed participants. In this framework, clients generate zero-knowledge proofs to demonstrate that local model updates are produced through legitimate training procedures without revealing the underlying data or model parameters. By integrating cryptographic verification with decentralized coordination, the proposed approach establishes trust among mutually untrusted participants while avoiding the computational cost of traditional blockchain consensus mechanisms. Second, it introduces a Layer-wise Optimal Transport-based Detection (LOTD) framework to identify malicious client updates in federated fine-tuning of large models using Low-Rank Adaptation (LoRA). The proposed method analyzes distributional changes in LoRA parameters through the Wasserstein distance and aggregates layer-wise discrepancies to detect abnormal updates. This design enables the detection of backdoor attacks that are difficult to identify using conventional parameter-distance methods and improves the robustness of federated model adaptation. Third, it develops Behavioral Hard Probability Optimization (BHPO), a preference-based optimization method designed to improve the robustness of LLMs against prompt injection attacks. Unlike conventional preference optimization methods that rely primarily on margin-based objectives, BHPO directly shapes the output probability distribution to suppress unsafe responses while preserving the likelihood of safe behaviors. This approach strengthens the model’s resistance to adversarial instructions while maintaining overall task performance. Extensive experiments on multiple datasets and attack settings demonstrate that the proposed methods effectively improve the security and reliability of AI systems during training and deployment. In summary, the principal contributions of this thesis are: • Verifiable federated learning: A blockchain-based federated learning framework with Zero-Knowledge Proof of Training (ZKPoT), enabling participants to verify legitimate local training without exposing private data or model parameters. • Robust federated fine-tuning: A Layer-wise Optimal Transport-based Detection (LOTD) framework for identifying and filtering malicious LoRA updates, thereby mitigating backdoor attacks while preserving clean-task performance. • Reliable LLM deployment: Behavioral Hard Probability Optimization (BHPO), a preference-based optimization method that improves resistance to prompt injection while maintaining utility on benign tasks.<p></p>
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Tianxing Fu (21044168)
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- All rights reserved
Identifiers
dc:identifier.*- Identifier
- 10779/exe.33085994.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/33085994