National University of Singapore
TOWARDS RELIABLE AI UNDER DISTRIBUTION SHIFTS: A DATA-CENTRIC PERSPECTIVE
Abstract
dc:description.abstractMachine learning (ML) models often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we closely examine the effect of low-quality data on model generalization. We theoretically show that when training with empirical risk minimization, label noise exacerbates the effect of spurious correlations in the training data. Second, we introduce two data-centric strategies to diagnose and improve quality of data. To detect mislabeled data, we propose an efficient approximation for the Shapley value. To improve the quality of data, we develop an approach that reweights the training data to enhance the robustness to subpopulation shift. Finally, we address the shortage of high-quality data. Specifically, we devise a game-theoretic framework, collaborative causal inference, which incentivizes data sharing among self-interested parties for causal inference.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- QIAO RUI