Back to results

University of Illinois at Urbana-Champaign

Leveraging distributional context for safe and interactive data science at scale

Abstract

dc:description

"Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew ""shortcuts"" that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this ""recipe"" to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are ""tighter"", allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice."

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Macke, Stephen Thomas
Contributors dc:contributor
  • Parameswaran, Aditya
  • Sundaram, Hari
  • Tong, Hanghang
  • Beutel, Alex

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2021 Stephen Thomas Macke
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/113024
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/113024

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Macke, Stephen Thomas. Leveraging distributional context for safe and interactive data science at scale. Dissertation thesis, University of Illinois at Urbana-Champaign, 2022. http://hdl.handle.net/2142/113024