University of Missouri--Columbia
Communication-efficient and privacy-preserving statistical learning with applications in high dimensional data
Abstract
dc:description.abstract[EMBARGOED UNTIL 12/01/2026] This dissertation explores advanced methods within federated learning (FL), a paradigm that fundamentally shifts from traditional, centralized machine learning. Instead of aggregating raw data--which creates significant privacy vulnerabilities and complicates compliance with regulations like GDPR and HIPAA--federated learning brings the model to the data, enabling collaborative training while data remains permanently decentralized. Most existing FL methods successfully utilize deterministic, frequentist approaches for efficient model training. This work seeks to extend those frameworks by explicitly addressing the inherent heterogeneity across data sites and integrating robust mechanisms for quantifying uncertainty. To address this gap, this dissertation argues for the integration of a statistical, and specifically Bayesian, perspective. Unlike deterministic methods, a Bayesian framework provides a probabilistic approach that can incorporate prior knowledge, inherently model heterogeneity, and provide a comprehensive quantification of model uncertainty. This is particularly crucial for high-stakes domains like biomedical research, where understanding the confidence of a prediction is as important as the prediction itself. This research, therefore, focuses on the critical FL challenges of data heterogeneity, communication efficiency, and high-dimensional variable selection, proposing novel non-linear and Bayesian approaches to enhance the robustness, privacy, and statistical power of federated systems. Chapter 2 addresses the critical challenge of data heterogeneity. It introduces a novel inferential ”gatekeeping” framework to determine tolerable levels of heterogeneity before federated inference becomes unreliable. Building on this, two new algorithms, HSD-ODAL and HSD-ODAL-MEM, are proposed to automatically detect and separate homogeneous (globally shared) and heterogeneous (site-specific) model parameters, demonstrating superior performance over existing methods in handling heterogeneous data. Chapter 3 presents two new Bayesian frameworks for communication-efficient clustering of high-dimensional data. FLamb, an iterative server-based algorithm, and OSCLamb, a decentralized single-round framework, both extend the Latent Mixture (Lamb) model for distributed settings. Real-world applications demonstrate a key trade-off: FLamb provides a superior model fit, while OSCLamb offers significantly faster, more robust performance in settings with severe data heterogeneity. Chapter 4 tackles communication-efficient feature selection in Vertical Federated Learning (VFL) for unstructured data, such as images. The proposed Trans-VFL framework integrates transfer learning, using a frozen pre-trained model to create high-level feature vectors. A novel ”selection layer” is then pruned using a federated group lasso adaptation, enabling effective and efficient feature selection on informative learned embeddings rather than raw inputs. Collectively, this dissertation contributes a suite of novel algorithms and theoretical frameworks that significantly advance the practical capabilities of federated learning, enabling more sophisticated and statistically sound analysis of decentralized, high-dimensional, and heterogeneous data while preserving privacy.
Degree
thesis:*- Name thesis:degree_name
- Ph. D.
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Statistics
- Grantor dc:publisher
- University of Missouri--Columbia
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Huang, Yilun
- Advisor dc:contributor.advisor
-
- Chakraborty, Sounak
Rights
- Language dc:language.iso
- eng, English
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:mospace.umsystem.edu:10355/111070