{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/316924"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/316924","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"TOWARDS RELIABLE AI UNDER DISTRIBUTION SHIFTS: A DATA-CENTRIC PERSPECTIVE","abstract":"Machine learning (ML) models often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we closely examine the effect of low-quality data on model generalization. We theoretically show that when training with empirical risk minimization, label noise exacerbates the effect of spurious correlations in the training data. Second, we introduce two data-centric strategies to diagnose and improve quality of data. To detect mislabeled data, we propose an efficient approximation for the Shapley value. To improve the quality of data, we develop an approach that reweights the training data to enhance the robustness to subpopulation shift. Finally, we address the shortage of high-quality data. Specifically, we devise a game-theoretic framework, collaborative causal inference, which incentivizes data sharing among self-interested parties for causal inference.","abstract_html":"Machine learning (ML) models often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we closely examine the effect of low-quality data on model generalization. We theoretically show that when training with empirical risk minimization, label noise exacerbates the effect of spurious correlations in the training data. Second, we introduce two data-centric strategies to diagnose and improve quality of data. To detect mislabeled data, we propose an efficient approximation for the Shapley value. To improve the quality of data, we develop an approach that reweights the training data to enhance the robustness to subpopulation shift. Finally, we address the shortage of high-quality data. Specifically, we devise a game-theoretic framework, collaborative causal inference, which incentivizes data sharing among self-interested parties for causal inference.","abstract_has_math":false,"creators":["QIAO RUI"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-11","date_published":"2025-04-11","updated_at":"2026-07-24T03:31:13Z","subjects":["mislabeled data","data attribution","data-centric AI","spurious correlation","distribution shifts","Trustworthy and reliable AI"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/fed73e76-f0df-48cf-8fbb-4aea830f97c5/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["QIAO RUI"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-04-11"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/316924"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["mislabeled data","data attribution","data-centric AI","spurious correlation","distribution shifts","Trustworthy and reliable AI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/fed73e76-f0df-48cf-8fbb-4aea830f97c5/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/6377726b-232c-4ff5-8510-56435cc7db09/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Machine learning (ML) models often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we closely examine the effect of low-quality data on model generalization. We theoretically show that when training with empirical risk minimization, label noise exacerbates the effect of spurious correlations in the training data. Second, we introduce two data-centric strategies to diagnose and improve quality of data. To detect mislabeled data, we propose an efficient approximation for the Shapley value. To improve the quality of data, we develop an approach that reweights the training data to enhance the robustness to subpopulation shift. Finally, we address the shortage of high-quality data. Specifically, we devise a game-theoretic framework, collaborative causal inference, which incentivizes data sharing among self-interested parties for causal inference."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","67101b5903a2e5daec3f35437dcfe90e","6c0cb142bf6b03b9b891d0aa68a108f3"]},{"key":"dc:title","label":"Title","values":["TOWARDS RELIABLE AI UNDER DISTRIBUTION SHIFTS: A DATA-CENTRIC PERSPECTIVE"]}]}],"canonical_facts":{"dc:creator":["QIAO RUI"],"dc:date.issued":["2025-04-11"],"dc:description.abstract":["Machine learning (ML) models often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we closely examine the effect of low-quality data on model generalization. We theoretically show that when training with empirical risk minimization, label noise exacerbates the effect of spurious correlations in the training data. Second, we introduce two data-centric strategies to diagnose and improve quality of data. To detect mislabeled data, we propose an efficient approximation for the Shapley value. To improve the quality of data, we develop an approach that reweights the training data to enhance the robustness to subpopulation shift. Finally, we address the shortage of high-quality data. Specifically, we devise a game-theoretic framework, collaborative causal inference, which incentivizes data sharing among self-interested parties for causal inference."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","67101b5903a2e5daec3f35437dcfe90e","6c0cb142bf6b03b9b891d0aa68a108f3"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/6377726b-232c-4ff5-8510-56435cc7db09/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/316924"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/fed73e76-f0df-48cf-8fbb-4aea830f97c5/download"],"dc:subject":["mislabeled data","data attribution","data-centric AI","spurious correlation","distribution shifts","Trustworthy and reliable AI"],"dc:title":["TOWARDS RELIABLE AI UNDER DISTRIBUTION SHIFTS: A DATA-CENTRIC PERSPECTIVE"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:31:13Z"}