{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132484"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132484","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Heterogeneous machine learning with decentralized data","abstract":"This thesis explores heterogeneous machine learning with decentralized data, where multiple clients with distinct data distributions jointly train or adapt machine learning models under the coordination of a central server. Throughout the process, clients’ private data never leave their local devices. This paradigm underlies numerous real-world applications, such as collaborative training of financial fraud detection models among banks, or collective health monitoring enabled by massive wearable devices. We investigate three fundamental challenges in this setting. (P1) Effective model training: How can we train models from multiple source clients with heterogeneous labeled data so that models perform well across all clients? (P2) Adaptive model deployment: How can we adapt the trained model to each target client without labels, allowing it to adjust to its own data distribution for improved performance? (P3) Robust system design: How can we ensure robustness during both training and adaptation, preventing performance degradation caused by random failures or malicious attacks? To address (P1), we develop client clustering algorithms that enable knowledge transfer among clients with similar data distributions, allowing those with limited data to benefit from collaboration. For (P2), we first design task-specific adaptation algorithms that identify and mitigate distribution shifts across different modalities under one-to-one adaptation. We further extend to multi-client collaboration, where the model learns patterns of distribution shifts across clients to enable collaborative test-time adaptation. Finally, for (P3), we propose robust training algorithms resilient to abnormal or adversarial clients, and robust adaptation algorithms that mitigate model prediction bias and out-of-distribution data effects. Together, these contributions form an effective and robust framework for heterogeneous machine learning under decentralized data environments.","abstract_html":"This thesis explores heterogeneous machine learning with decentralized data, where multiple clients with distinct data distributions jointly train or adapt machine learning models under the coordination of a central server. Throughout the process, clients’ private data never leave their local devices. This paradigm underlies numerous real-world applications, such as collaborative training of financial fraud detection models among banks, or collective health monitoring enabled by massive wearable devices. We investigate three fundamental challenges in this setting. (P1) Effective model training: How can we train models from multiple source clients with heterogeneous labeled data so that models perform well across all clients? (P2) Adaptive model deployment: How can we adapt the trained model to each target client without labels, allowing it to adjust to its own data distribution for improved performance? (P3) Robust system design: How can we ensure robustness during both training and adaptation, preventing performance degradation caused by random failures or malicious attacks? To address (P1), we develop client clustering algorithms that enable knowledge transfer among clients with similar data distributions, allowing those with limited data to benefit from collaboration. For (P2), we first design task-specific adaptation algorithms that identify and mitigate distribution shifts across different modalities under one-to-one adaptation. We further extend to multi-client collaboration, where the model learns patterns of distribution shifts across clients to enable collaborative test-time adaptation. Finally, for (P3), we propose robust training algorithms resilient to abnormal or adversarial clients, and robust adaptation algorithms that mitigate model prediction bias and out-of-distribution data effects. Together, these contributions form an effective and robust framework for heterogeneous machine learning under decentralized data environments.","abstract_has_math":false,"creators":["Bao, Wenxuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["He, Jingrui","Zhang, Tong","Zhao, Han","Li, Pan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Distribution shift","Federated learning","Test-time adaptation"],"languages":["en"],"rights":["Copyright 2025 Wenxuan Bao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132484","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["He, Jingrui","Zhang, Tong","Zhao, Han","Li, Pan"]},{"key":"dc:creator","label":"Author","values":["Bao, Wenxuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Distribution shift","Federated learning","Test-time adaptation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Wenxuan Bao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132484"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis explores heterogeneous machine learning with decentralized data, where multiple clients with distinct data distributions jointly train or adapt machine learning models under the coordination of a central server. Throughout the process, clients’ private data never leave their local devices. This paradigm underlies numerous real-world applications, such as collaborative training of financial fraud detection models among banks, or collective health monitoring enabled by massive wearable devices. We investigate three fundamental challenges in this setting. (P1) Effective model training: How can we train models from multiple source clients with heterogeneous labeled data so that models perform well across all clients? (P2) Adaptive model deployment: How can we adapt the trained model to each target client without labels, allowing it to adjust to its own data distribution for improved performance? (P3) Robust system design: How can we ensure robustness during both training and adaptation, preventing performance degradation caused by random failures or malicious attacks? To address (P1), we develop client clustering algorithms that enable knowledge transfer among clients with similar data distributions, allowing those with limited data to benefit from collaboration. For (P2), we first design task-specific adaptation algorithms that identify and mitigate distribution shifts across different modalities under one-to-one adaptation. We further extend to multi-client collaboration, where the model learns patterns of distribution shifts across clients to enable collaborative test-time adaptation. Finally, for (P3), we propose robust training algorithms resilient to abnormal or adversarial clients, and robust adaptation algorithms that mitigate model prediction bias and out-of-distribution data effects. Together, these contributions form an effective and robust framework for heterogeneous machine learning under decentralized data environments.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Wenxuan Bao, accepted the attached license on 2025-11-11 at 16:46.","The student, Wenxuan Bao, submitted this Dissertation for approval on 2025-11-11 at 16:51.","This Dissertation was approved for publication on 2025-11-12 at 11:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22858 on 2026-02-19 at 18:24:34"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Heterogeneous machine learning with decentralized data"]}]}],"canonical_facts":{"dc:contributor":["He, Jingrui","Zhang, Tong","Zhao, Han","Li, Pan"],"dc:creator":["Bao, Wenxuan"],"dc:date":["2025-12","2025-11-12"],"dc:description":["This thesis explores heterogeneous machine learning with decentralized data, where multiple clients with distinct data distributions jointly train or adapt machine learning models under the coordination of a central server. Throughout the process, clients’ private data never leave their local devices. This paradigm underlies numerous real-world applications, such as collaborative training of financial fraud detection models among banks, or collective health monitoring enabled by massive wearable devices. We investigate three fundamental challenges in this setting. (P1) Effective model training: How can we train models from multiple source clients with heterogeneous labeled data so that models perform well across all clients? (P2) Adaptive model deployment: How can we adapt the trained model to each target client without labels, allowing it to adjust to its own data distribution for improved performance? (P3) Robust system design: How can we ensure robustness during both training and adaptation, preventing performance degradation caused by random failures or malicious attacks? To address (P1), we develop client clustering algorithms that enable knowledge transfer among clients with similar data distributions, allowing those with limited data to benefit from collaboration. For (P2), we first design task-specific adaptation algorithms that identify and mitigate distribution shifts across different modalities under one-to-one adaptation. We further extend to multi-client collaboration, where the model learns patterns of distribution shifts across clients to enable collaborative test-time adaptation. Finally, for (P3), we propose robust training algorithms resilient to abnormal or adversarial clients, and robust adaptation algorithms that mitigate model prediction bias and out-of-distribution data effects. Together, these contributions form an effective and robust framework for heterogeneous machine learning under decentralized data environments.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Wenxuan Bao, accepted the attached license on 2025-11-11 at 16:46.","The student, Wenxuan Bao, submitted this Dissertation for approval on 2025-11-11 at 16:51.","This Dissertation was approved for publication on 2025-11-12 at 11:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22858 on 2026-02-19 at 18:24:34"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132484"],"dc:language":["en"],"dc:rights":["Copyright 2025 Wenxuan Bao"],"dc:subject":["Distribution shift","Federated learning","Test-time adaptation"],"dc:title":["Heterogeneous machine learning with decentralized data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}