{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127396"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127396","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Invariant learning for learning in the wild","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-12-01","abstract_has_math":false,"creators":["Wang, Xiaoyang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Koyejo, Oluwasanmi","Nahrstedt, Klara","Tong, Hanghang","Dimitriadis, Dimitrios"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-04","date_published":"2024-12-04","updated_at":"2026-07-22T22:25:04Z","subjects":["Machine Learning","Invariance","Robustness"],"languages":["en","eng"],"rights":["Copyright 2024 Xiaoyang Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127396","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Koyejo, Oluwasanmi","Nahrstedt, Klara","Tong, Hanghang","Dimitriadis, Dimitrios"]},{"key":"dc:creator","label":"Author","values":["Wang, Xiaoyang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-12-04","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Invariance","Robustness"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Xiaoyang Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127396"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Xiaoyang Wang, accepted the attached license on 2024-12-04 at 02:22.","The student, Xiaoyang Wang, submitted this Dissertation for approval on 2024-12-04 at 02:30.","This Dissertation was approved for publication on 2024-12-04 at 13:32.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21491 on 2025-03-28 at 14:44:40","Machine learning models are increasingly deployed in production (i.e., the wild) but may fail for various reasons. For example, fraud detection models can protect numerous users against phishing emails but are subject to intentional poisoning and may fail to identify novel types of phishing. Similarly, a language model provides timely answers to user questions. However, the answer quality can decrease significantly or be harmful even if minor changes apply to the questions. Common failures of machine learning models in production environments fall into two categories: (1) data quality and (2) data shift. Data quality problems can be caused by malicious adversaries that aim to corrupt machine learning models, uncurated crowdsourced data from the web, etc. Meanwhile, data shift problems often occur due to the mismatch between the offline training data and the continuously evolving data in online production environments. Tackling the data quality and shift problems requires methods that help machine learning models continuously learn generally useful patterns from the data without entangling the harmful ones. In this dissertation, we introduce invariant learning as a paradigm to meet the aforementioned requirement and address the data quality and shift problems in the wild. In particular, we first study a data quality problem with multiple data sources with mixed data qualities. Our main contribution to this problem is a novel algorithm that helps machine learning models learn invariant patterns from multiple data sources and selectively filter out the contribution of low-quality data. Then, we further study a setting that requires machine learning models to be fine-tuned (i.e., customized) to a particular data source with improved performance but does not sacrifice the invariance benefit. The last part of this dissertation applies invariant learning to an active fine-tuning problem, which requires machine learning models to continuously learn new data with improved data efficiency. Our invariance-aware approach selects subsets of data samples that invariantly benefit the full dataset with minimal neglect of unselected data samples and helps machine learning models adapt to shifting data more effectively."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Invariant learning for learning in the wild"]}]}],"canonical_facts":{"dc:contributor":["Koyejo, Oluwasanmi","Nahrstedt, Klara","Tong, Hanghang","Dimitriadis, Dimitrios"],"dc:creator":["Wang, Xiaoyang"],"dc:date":["2024-12-04","2024-12"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Xiaoyang Wang, accepted the attached license on 2024-12-04 at 02:22.","The student, Xiaoyang Wang, submitted this Dissertation for approval on 2024-12-04 at 02:30.","This Dissertation was approved for publication on 2024-12-04 at 13:32.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21491 on 2025-03-28 at 14:44:40","Machine learning models are increasingly deployed in production (i.e., the wild) but may fail for various reasons. For example, fraud detection models can protect numerous users against phishing emails but are subject to intentional poisoning and may fail to identify novel types of phishing. Similarly, a language model provides timely answers to user questions. However, the answer quality can decrease significantly or be harmful even if minor changes apply to the questions. Common failures of machine learning models in production environments fall into two categories: (1) data quality and (2) data shift. Data quality problems can be caused by malicious adversaries that aim to corrupt machine learning models, uncurated crowdsourced data from the web, etc. Meanwhile, data shift problems often occur due to the mismatch between the offline training data and the continuously evolving data in online production environments. Tackling the data quality and shift problems requires methods that help machine learning models continuously learn generally useful patterns from the data without entangling the harmful ones. In this dissertation, we introduce invariant learning as a paradigm to meet the aforementioned requirement and address the data quality and shift problems in the wild. In particular, we first study a data quality problem with multiple data sources with mixed data qualities. Our main contribution to this problem is a novel algorithm that helps machine learning models learn invariant patterns from multiple data sources and selectively filter out the contribution of low-quality data. Then, we further study a setting that requires machine learning models to be fine-tuned (i.e., customized) to a particular data source with improved performance but does not sacrifice the invariance benefit. The last part of this dissertation applies invariant learning to an active fine-tuning problem, which requires machine learning models to continuously learn new data with improved data efficiency. Our invariance-aware approach selects subsets of data samples that invariantly benefit the full dataset with minimal neglect of unselected data samples and helps machine learning models adapt to shifting data more effectively."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127396"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Xiaoyang Wang"],"dc:subject":["Machine Learning","Invariance","Robustness"],"dc:title":["Invariant learning for learning in the wild"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}