University of Illinois at Urbana-Champaign
Leveraging distributional context for safe and interactive data science at scale
Abstract
dc:description"Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew ""shortcuts"" that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this ""recipe"" to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are ""tighter"", allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice."
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Macke, Stephen Thomas
- Contributors dc:contributor
-
- Parameswaran, Aditya
- Sundaram, Hari
- Tong, Hanghang
- Beutel, Alex
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2021 Stephen Thomas Macke
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/113024
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/113024