{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113024"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113024","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Leveraging distributional context for safe and interactive data science at scale","abstract":"\"Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew \"\"shortcuts\"\" that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this \"\"recipe\"\" to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are \"\"tighter\"\", allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice.\"","abstract_html":"&quot;Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew &quot;&quot;shortcuts&quot;&quot; that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this &quot;&quot;recipe&quot;&quot; to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are &quot;&quot;tighter&quot;&quot;, allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice.&quot;","abstract_has_math":false,"creators":["Macke, Stephen Thomas"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Parameswaran, Aditya","Sundaram, Hari","Tong, Hanghang","Beutel, Alex"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:39Z","date_published":"2022-01-12T21:45:39Z","updated_at":"2026-07-22T22:24:52Z","subjects":["data science","visualization","EDA","computational notebooks","approximate query processing"],"languages":["en"],"rights":["Copyright 2021 Stephen Thomas Macke"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113024","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Parameswaran, Aditya","Sundaram, Hari","Tong, Hanghang","Beutel, Alex"]},{"key":"dc:creator","label":"Author","values":["Macke, Stephen Thomas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:39Z","2021-07-14","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["data science","visualization","EDA","computational notebooks","approximate query processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Stephen Thomas Macke"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113024"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew \"\"shortcuts\"\" that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this \"\"recipe\"\" to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are \"\"tighter\"\", allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Stephen Macke, accepted the attached license on 2021-07-12 at 14:10.","The student, Stephen Macke, submitted this Dissertation for approval on 2021-07-12 at 14:24.","This Dissertation was approved for publication on 2021-07-14 at 14:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16871 on 2022-01-12 at 12:45:02","Made available in DSpace on 2022-01-12T21:45:39Z (GMT). No. of bitstreams: 3 MACKE-DISSERTATION-2021.pdf: 4672605 bytes, checksum: 27a0d1f6c54885cd0f0b987dd3f0660e (MD5) LICENSE.txt: 4210 bytes, checksum: e5ffd22faabe3c5df448e28d9f194cec (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 6dd42d048a3198667675542db7c1532e (MD5) Previous issue date: 2021-07-14"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Leveraging distributional context for safe and interactive data science at scale"]}]}],"canonical_facts":{"dc:contributor":["Parameswaran, Aditya","Sundaram, Hari","Tong, Hanghang","Beutel, Alex"],"dc:creator":["Macke, Stephen Thomas"],"dc:date":["2022-01-12T21:45:39Z","2021-07-14","2021-08"],"dc:description":["\"Data science is an iterative, exploratory, and ad-hoc process performed by individuals and teams possessing increasingly varied backgrounds and skill-sets. As such, we need data science to be interactive, so that data scientists are not bottlenecked when trying out new hypotheses or confirming existing ones. Moreover, data science must be safe, ensuring that data scientists, especially those with limited programming or analysis experience, avoid making incorrect inferences. Safety and interactivity are typically at odds with one another, since various notions of safety often eschew \"\"shortcuts\"\" that make working with large-scale data tractable. In this dissertation, we aim to meet the dual objectives of interactivity and safety at scale by leveraging distributional context—specifically the distributions of the data and the operations performed by data scientists. We apply this \"\"recipe\"\" to five different key data science settings: (i) for machine learning development, we provide context-aware caching algorithms that allow model developers to benefit from interactive iteration times during model development, while not requiring error-prone manual tracking of reusable intermediates; (ii) for visualization search, we develop context-aware sampling algorithms that support interactive search for patterns in visualizations, while ensuring that the results meet rigorous quality guarantees; (iii) for browsing, we develop workload-aware learned Bloom filters optimized for multidimensional data that allow analysts to quickly identify records that have been examined before, all while guarding against false negatives; (iv) for report generation, we develop context-aware aggregate approximation algorithms that provide rigorous distribution-aware confidence intervals around aggregates, while ensuring that the intervals are \"\"tighter\"\", allowing analysts to make decisions sooner; and (v) finally, for error-prone interactions in computational notebooks, we demonstrate approximate lineage-capture techniques that warn data scientists of unsafe cell executions for many cases encountered in practice.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Stephen Macke, accepted the attached license on 2021-07-12 at 14:10.","The student, Stephen Macke, submitted this Dissertation for approval on 2021-07-12 at 14:24.","This Dissertation was approved for publication on 2021-07-14 at 14:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16871 on 2022-01-12 at 12:45:02","Made available in DSpace on 2022-01-12T21:45:39Z (GMT). No. of bitstreams: 3 MACKE-DISSERTATION-2021.pdf: 4672605 bytes, checksum: 27a0d1f6c54885cd0f0b987dd3f0660e (MD5) LICENSE.txt: 4210 bytes, checksum: e5ffd22faabe3c5df448e28d9f194cec (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 6dd42d048a3198667675542db7c1532e (MD5) Previous issue date: 2021-07-14"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113024"],"dc:language":["en"],"dc:rights":["Copyright 2021 Stephen Thomas Macke"],"dc:subject":["data science","visualization","EDA","computational notebooks","approximate query processing"],"dc:title":["Leveraging distributional context for safe and interactive data science at scale"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}