University of Illinois - Chicago
Data-centric Approaches for Responsible Data Science
Abstract
dc:descriptionThe abundance of data, coupled with recent advancements in computation, has revolutionized almost every aspect of human life. While the undeniable benefits of this evolution are evident and despite the promise to bring good to human life and society, data-driven technologies could instead become harmful if not used responsibly. These harms usually stem from the inherent biases within the data and ``an algorithm is only as good as the data it works with''; therefore, if not addressed promptly, biases could get amplified to the downstream tasks in the data science pipelines. Without addressing the bias issues, we cannot expect AI-based societal solutions to have equitable outcomes. To this end, the main theme of this study revolves around responsible data science and algorithmic fairness with a strong emphasis on data-centric approaches. In this study, we firstly focus on data coverage as a data-centric approach for identifying and resolving the misrepresentation of minorities in data. We propose novel algorithms that identify insufficient data coverage across data with different modalities and use a lack of representation information to generate data-centric reliability warnings. Secondly, we study entity matching—a foundational task in data integration—through the lens of group fairness. We introduce new fairness definitions for entity matching, conduct a broad empirical analysis of existing techniques, and propose efficient algorithms to identify unbiased datasets within large search spaces. Building on this, we develop a framework for auditing entity matchers, diagnosing underlying reasons of unfairness, and resolving the issues through human-in-the-loop exploration with an ensemble of matchers. Thirdly, we revisit several classic algorithmic data structures and problems through the lens of group fairness. We propose FairHash, a data-dependent hashmap that ensures group-level uniform distribution across buckets and simultaneously satisfies three formal fairness notions. For frequency estimation, we introduce Fair-Count-Min, a sketch that guarantees equal approximation factors across groups using group-aware semi-uniform hashing. We also revisit the Weighted Set Multi-Cover problem, a fundamental problem with significant applications in group fairness and package recommendations. We show that when the size of the universe is bounded, improved approximation algorithms can be designed. Ultimately, this dissertation highlights the pivotal role of data in achieving equitable outcomes and establishes a foundation for future work in responsible data science.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nima Shahbazi (11048343)
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- In Copyright
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.31451296.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/31451296