Back to results

University of Illinois at Urbana-Champaign

Invariant learning for learning in the wild

Abstract

dc:description

Machine learning models are increasingly deployed in production (i.e., the wild) but may fail for various reasons. For example, fraud detection models can protect numerous users against phishing emails but are subject to intentional poisoning and may fail to identify novel types of phishing. Similarly, a language model provides timely answers to user questions. However, the answer quality can decrease significantly or be harmful even if minor changes apply to the questions. Common failures of machine learning models in production environments fall into two categories: (1) data quality and (2) data shift. Data quality problems can be caused by malicious adversaries that aim to corrupt machine learning models, uncurated crowdsourced data from the web, etc. Meanwhile, data shift problems often occur due to the mismatch between the offline training data and the continuously evolving data in online production environments. Tackling the data quality and shift problems requires methods that help machine learning models continuously learn generally useful patterns from the data without entangling the harmful ones. In this dissertation, we introduce invariant learning as a paradigm to meet the aforementioned requirement and address the data quality and shift problems in the wild. In particular, we first study a data quality problem with multiple data sources with mixed data qualities. Our main contribution to this problem is a novel algorithm that helps machine learning models learn invariant patterns from multiple data sources and selectively filter out the contribution of low-quality data. Then, we further study a setting that requires machine learning models to be fine-tuned (i.e., customized) to a particular data source with improved performance but does not sacrifice the invariance benefit. The last part of this dissertation applies invariant learning to an active fine-tuning problem, which requires machine learning models to continuously learn new data with improved data efficiency. Our invariance-aware approach selects subsets of data samples that invariantly benefit the full dataset with minimal neglect of unselected data samples and helps machine learning models adapt to shifting data more effectively.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wang, Xiaoyang
Contributors dc:contributor
  • Koyejo, Oluwasanmi
  • Nahrstedt, Klara
  • Tong, Hanghang
  • Dimitriadis, Dimitrios

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Xiaoyang Wang
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127396

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Wang, Xiaoyang. Invariant learning for learning in the wild. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127396