{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129820"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129820","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Heterogeneous machine learning: characterization, generation and comprehension","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Zheng, Lecheng"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["He, JIngrui","Sun, Jimeng","Wang, Yuxiong","Birge, John"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-06-16","date_published":"2025-06-16","updated_at":"2026-07-22T22:25:05Z","subjects":["Heterogeneous Learning","Multi-view Learning","Multi-label Learning","Contrastive Learning","Self-supervised Learning"],"languages":["en","eng"],"rights":["Copyright 2025 Lecheng Zheng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129820","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["He, JIngrui","Sun, Jimeng","Wang, Yuxiong","Birge, John"]},{"key":"dc:creator","label":"Author","values":["Zheng, Lecheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-06-16","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Heterogeneous Learning","Multi-view Learning","Multi-label Learning","Contrastive Learning","Self-supervised Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Lecheng Zheng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129820"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Lecheng Zheng, accepted the attached license on 2025-06-13 at 13:54.","The student, Lecheng Zheng, submitted this Dissertation for approval on 2025-06-13 at 14:02.","This Dissertation was approved for publication on 2025-06-16 at 10:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22333 on 2025-10-20 at 16:57:05","In the era of big data, the increasing complexity of real-world applications has brought data and model heterogeneity to the forefront of machine learning research. Domains such as sustainability, finance, and agriculture rely on the integration of diverse data sources, including climate time series (e.g., MERRA2, ERA5) and satellite imagery (e.g., Landsat8, SAR), to support tasks ranging from weather forecasting and extreme event analysis to broader climate impact assessments. These tasks highlight the inherent data heterogeneity across modalities and objectives, while simultaneously motivating the need for model heterogeneity to address varying learning paradigms. The recent surge in foundation models, such as large language models, vision transformers, and graph neural networks, has further accentuated the challenge of unifying diverse model architectures, thereby giving rise to a new research paradigm: heterogeneous machine learning (HML). HML aims to improve generalizability, adaptability, and trustworthiness by leveraging the richness of data and model diversity. Despite its promise, heterogeneous machine learning poses significant challenges, including characterizing complex heterogeneity involving multiple modalities (view heterogeneity) and task objectives (label heterogeneity), ensuring trustworthy machine learning that adheres to principles of privacy, fairness, robustness, and interpretability, handling incomplete information, a common issue in large-scale heterogeneous datasets where missing views or partial labels are prevalent and establishing strong theoretical foundations to guide the design, analysis, and interpretation of heterogeneous models. Addressing these challenges requires a systematic and interdisciplinary approach. In this thesis, my research tackles these challenges through three interconnected pillars, i.e., Heterogeneous Machine Learning Characterization, Heterogeneous Machine Learning Generation, and Heterogeneous Machine Learning Comprehension. First, in Heterogeneous Machine Learning Characterization, I focus on understanding and modeling the structure of heterogeneous data by capturing inter-modality correlations and task relatedness. I also integrate principles of trustworthy AI to ensure fairness, interpretability, and robustness throughout the modeling process. Second, in Heterogeneous Machine Learning Generation, I develop generative models to address data incompleteness by learning plausible patterns and restoring missing modalities or labels. These models support applications such as data augmentation, anomaly detection, and secure learning in domains where data quality is often compromised. Finally, in Heterogeneous Machine Learning Comprehension, I aim to bridge the gap between empirical success and theoretical understanding. This includes deriving fairness bounds, leveraging information-theoretic frameworks for contrastive learning, and using graph theory and influence functions to analyze model behavior and detect anomalies."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Heterogeneous machine learning: characterization, generation and comprehension"]}]}],"canonical_facts":{"dc:contributor":["He, JIngrui","Sun, Jimeng","Wang, Yuxiong","Birge, John"],"dc:creator":["Zheng, Lecheng"],"dc:date":["2025-06-16","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Lecheng Zheng, accepted the attached license on 2025-06-13 at 13:54.","The student, Lecheng Zheng, submitted this Dissertation for approval on 2025-06-13 at 14:02.","This Dissertation was approved for publication on 2025-06-16 at 10:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22333 on 2025-10-20 at 16:57:05","In the era of big data, the increasing complexity of real-world applications has brought data and model heterogeneity to the forefront of machine learning research. Domains such as sustainability, finance, and agriculture rely on the integration of diverse data sources, including climate time series (e.g., MERRA2, ERA5) and satellite imagery (e.g., Landsat8, SAR), to support tasks ranging from weather forecasting and extreme event analysis to broader climate impact assessments. These tasks highlight the inherent data heterogeneity across modalities and objectives, while simultaneously motivating the need for model heterogeneity to address varying learning paradigms. The recent surge in foundation models, such as large language models, vision transformers, and graph neural networks, has further accentuated the challenge of unifying diverse model architectures, thereby giving rise to a new research paradigm: heterogeneous machine learning (HML). HML aims to improve generalizability, adaptability, and trustworthiness by leveraging the richness of data and model diversity. Despite its promise, heterogeneous machine learning poses significant challenges, including characterizing complex heterogeneity involving multiple modalities (view heterogeneity) and task objectives (label heterogeneity), ensuring trustworthy machine learning that adheres to principles of privacy, fairness, robustness, and interpretability, handling incomplete information, a common issue in large-scale heterogeneous datasets where missing views or partial labels are prevalent and establishing strong theoretical foundations to guide the design, analysis, and interpretation of heterogeneous models. Addressing these challenges requires a systematic and interdisciplinary approach. In this thesis, my research tackles these challenges through three interconnected pillars, i.e., Heterogeneous Machine Learning Characterization, Heterogeneous Machine Learning Generation, and Heterogeneous Machine Learning Comprehension. First, in Heterogeneous Machine Learning Characterization, I focus on understanding and modeling the structure of heterogeneous data by capturing inter-modality correlations and task relatedness. I also integrate principles of trustworthy AI to ensure fairness, interpretability, and robustness throughout the modeling process. Second, in Heterogeneous Machine Learning Generation, I develop generative models to address data incompleteness by learning plausible patterns and restoring missing modalities or labels. These models support applications such as data augmentation, anomaly detection, and secure learning in domains where data quality is often compromised. Finally, in Heterogeneous Machine Learning Comprehension, I aim to bridge the gap between empirical success and theoretical understanding. This includes deriving fairness bounds, leveraging information-theoretic frameworks for contrastive learning, and using graph theory and influence functions to analyze model behavior and detect anomalies."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129820"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Lecheng Zheng"],"dc:subject":["Heterogeneous Learning","Multi-view Learning","Multi-label Learning","Contrastive Learning","Self-supervised Learning"],"dc:title":["Heterogeneous machine learning: characterization, generation and comprehension"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}