{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105721"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105721","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Accelerating human-in-the-loop machine learning","abstract":"Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Unfortunately, existing machine learning systems focus narrowly on model training—a small fraction of the overall development time—and neglect to address iterative development. We propose Helix, a machine learning system that optimizes the execution across iterations—intelligently caching and reusing, or recomputing intermediates as appropriate. Helix captures a wide variety of application needs within its Scala DSL, with succinct syntax defining unified processes for data preprocessing, model specification, and learning. We demonstrate that the reuse problem can be cast as a Max-Flow problem, while the caching problem is NP-Hard. We develop effective lightweight heuristics for the latter. Empirical evaluation shows that Helix is not only able to handle a wide variety of use cases in one unified workflow but also much faster, providing run time reductions of up to 19× over state-of-the-art systems, such as DeepDive or KeystoneML, on four real-world applications in natural language processing, computer vision, social and natural sciences.","abstract_html":"Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Unfortunately, existing machine learning systems focus narrowly on model training—a small fraction of the overall development time—and neglect to address iterative development. We propose Helix, a machine learning system that optimizes the execution across iterations—intelligently caching and reusing, or recomputing intermediates as appropriate. Helix captures a wide variety of application needs within its Scala DSL, with succinct syntax defining unified processes for data preprocessing, model specification, and learning. We demonstrate that the reuse problem can be cast as a Max-Flow problem, while the caching problem is NP-Hard. We develop effective lightweight heuristics for the latter. Empirical evaluation shows that Helix is not only able to handle a wide variety of use cases in one unified workflow but also much faster, providing run time reductions of up to 19× over state-of-the-art systems, such as DeepDive or KeystoneML, on four real-world applications in natural language processing, computer vision, social and natural sciences.","abstract_has_math":false,"creators":["Xin, Doris"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Parameswaran, Aditya G"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:35:17Z","date_published":"2019-11-26T20:35:17Z","updated_at":"2026-07-22T22:24:44Z","subjects":["data management","machine learning","human-in-the-loop analytics"],"languages":["en"],"rights":["Copyright 2019 Doris Xin"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105721","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Parameswaran, Aditya G"]},{"key":"dc:creator","label":"Author","values":["Xin, Doris"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:35:17Z","2019-07-18","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["data management","machine learning","human-in-the-loop analytics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Doris Xin"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105721"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Unfortunately, existing machine learning systems focus narrowly on model training—a small fraction of the overall development time—and neglect to address iterative development. We propose Helix, a machine learning system that optimizes the execution across iterations—intelligently caching and reusing, or recomputing intermediates as appropriate. Helix captures a wide variety of application needs within its Scala DSL, with succinct syntax defining unified processes for data preprocessing, model specification, and learning. We demonstrate that the reuse problem can be cast as a Max-Flow problem, while the caching problem is NP-Hard. We develop effective lightweight heuristics for the latter. Empirical evaluation shows that Helix is not only able to handle a wide variety of use cases in one unified workflow but also much faster, providing run time reductions of up to 19× over state-of-the-art systems, such as DeepDive or KeystoneML, on four real-world applications in natural language processing, computer vision, social and natural sciences.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Doris Xin, accepted the attached license on 2019-07-18 at 12:43.","The student, Doris Xin, submitted this Thesis for approval on 2019-07-18 at 12:54.","This Thesis was approved for publication on 2019-07-18 at 16:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14380 on 2019-11-26 at 12:54:19","Made available in DSpace on 2019-11-26T20:35:17Z (GMT). No. of bitstreams: 3 XIN-THESIS-2019.pdf: 8362570 bytes, checksum: cc1f9c483ed6b7d54279c17863c50fb4 (MD5) DorisXinMSThesis.zip: 8221643 bytes, checksum: 00292aae6bc5111e3abc02a8a8521263 (MD5) LICENSE.txt: 4206 bytes, checksum: e95969723831a2a9a9d2b630630bd20f (MD5) Previous issue date: 2019-07-18"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Accelerating human-in-the-loop machine learning"]}]}],"canonical_facts":{"dc:contributor":["Parameswaran, Aditya G"],"dc:creator":["Xin, Doris"],"dc:date":["2019-11-26T20:35:17Z","2019-07-18","2019-08"],"dc:description":["Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Unfortunately, existing machine learning systems focus narrowly on model training—a small fraction of the overall development time—and neglect to address iterative development. We propose Helix, a machine learning system that optimizes the execution across iterations—intelligently caching and reusing, or recomputing intermediates as appropriate. Helix captures a wide variety of application needs within its Scala DSL, with succinct syntax defining unified processes for data preprocessing, model specification, and learning. We demonstrate that the reuse problem can be cast as a Max-Flow problem, while the caching problem is NP-Hard. We develop effective lightweight heuristics for the latter. Empirical evaluation shows that Helix is not only able to handle a wide variety of use cases in one unified workflow but also much faster, providing run time reductions of up to 19× over state-of-the-art systems, such as DeepDive or KeystoneML, on four real-world applications in natural language processing, computer vision, social and natural sciences.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Doris Xin, accepted the attached license on 2019-07-18 at 12:43.","The student, Doris Xin, submitted this Thesis for approval on 2019-07-18 at 12:54.","This Thesis was approved for publication on 2019-07-18 at 16:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14380 on 2019-11-26 at 12:54:19","Made available in DSpace on 2019-11-26T20:35:17Z (GMT). No. of bitstreams: 3 XIN-THESIS-2019.pdf: 8362570 bytes, checksum: cc1f9c483ed6b7d54279c17863c50fb4 (MD5) DorisXinMSThesis.zip: 8221643 bytes, checksum: 00292aae6bc5111e3abc02a8a8521263 (MD5) LICENSE.txt: 4206 bytes, checksum: e95969723831a2a9a9d2b630630bd20f (MD5) Previous issue date: 2019-07-18"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105721"],"dc:language":["en"],"dc:rights":["Copyright 2019 Doris Xin"],"dc:subject":["data management","machine learning","human-in-the-loop analytics"],"dc:title":["Accelerating human-in-the-loop machine learning"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}