{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97707"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97707","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Mitigating Spark straggler tasks for iterative applications by data re-partitioning","abstract":"Many of the data science applications nowadays feature large datasets and short tasks that run many iterations. When running these applications on a parallel processing framework like Apache Spark, one problem that affects the running time is the straggler, where a disproportionate long-running task slows down the entire cluster. In this work we present a straggler mitigation technique tailored for applications that run small tasks for many iterations over a large dataset, and implemented the algorithm in Apache Spark. We monitor the resources available on each Spark node, and dynamically re partition the dataset proportional to the estimated resource available. We have shown that our algorithm has negligible overhead for resource monitoring, and can improve the performance of Spark cluster significantly when stragglers are present.","abstract_html":"Many of the data science applications nowadays feature large datasets and short tasks that run many iterations. When running these applications on a parallel processing framework like Apache Spark, one problem that affects the running time is the straggler, where a disproportionate long-running task slows down the entire cluster. In this work we present a straggler mitigation technique tailored for applications that run small tasks for many iterations over a large dataset, and implemented the algorithm in Apache Spark. We monitor the resources available on each Spark node, and dynamically re partition the dataset proportional to the estimated resource available. We have shown that our algorithm has negligible overhead for resource monitoring, and can improve the performance of Spark cluster significantly when stragglers are present.","abstract_has_math":false,"creators":["Teng, Bo"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Campbell, Roy H."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T20:32:56Z","date_published":"2017-08-10T20:32:56Z","updated_at":"2026-07-22T22:24:34Z","subjects":["Straggler","Machine learning","Apache Spark","Iterative application"],"languages":["en"],"rights":["Copyright 2017 Bo Teng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97707","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Campbell, Roy H."]},{"key":"dc:creator","label":"Author","values":["Teng, Bo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T20:32:56Z","2019-08-11T09:15:24Z","2017-04-18","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Straggler","Machine learning","Apache Spark","Iterative application"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Bo Teng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97707"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Many of the data science applications nowadays feature large datasets and short tasks that run many iterations. When running these applications on a parallel processing framework like Apache Spark, one problem that affects the running time is the straggler, where a disproportionate long-running task slows down the entire cluster. In this work we present a straggler mitigation technique tailored for applications that run small tasks for many iterations over a large dataset, and implemented the algorithm in Apache Spark. We monitor the resources available on each Spark node, and dynamically re partition the dataset proportional to the estimated resource available. We have shown that our algorithm has negligible overhead for resource monitoring, and can improve the performance of Spark cluster significantly when stragglers are present.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Bo Teng, accepted the attached license on 2017-04-14 at 16:17.","The student, Bo Teng, submitted this Thesis for approval on 2017-04-14 at 16:27.","This Thesis was approved for publication on 2017-04-18 at 11:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10768 on 2017-08-10 at 15:05:37","Made available in DSpace on 2017-08-10T20:32:56Z (GMT). No. of bitstreams: 2 TENG-THESIS-2017.pdf: 486378 bytes, checksum: 7806fa81f8e97de0ae61fb5e33e81115 (MD5) LICENSE.txt: 4204 bytes, checksum: 8a65a459f7723bf11d5082e14db4b63a (MD5) Previous issue date: 2017-04-18","Embargo set by: Colleen Fallaw for item 102760 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102760 on 2019-08-11T09:15:24Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Mitigating Spark straggler tasks for iterative applications by data re-partitioning"]}]}],"canonical_facts":{"dc:contributor":["Campbell, Roy H."],"dc:creator":["Teng, Bo"],"dc:date":["2017-08-10T20:32:56Z","2019-08-11T09:15:24Z","2017-04-18","2017-05"],"dc:description":["Many of the data science applications nowadays feature large datasets and short tasks that run many iterations. When running these applications on a parallel processing framework like Apache Spark, one problem that affects the running time is the straggler, where a disproportionate long-running task slows down the entire cluster. In this work we present a straggler mitigation technique tailored for applications that run small tasks for many iterations over a large dataset, and implemented the algorithm in Apache Spark. We monitor the resources available on each Spark node, and dynamically re partition the dataset proportional to the estimated resource available. We have shown that our algorithm has negligible overhead for resource monitoring, and can improve the performance of Spark cluster significantly when stragglers are present.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Bo Teng, accepted the attached license on 2017-04-14 at 16:17.","The student, Bo Teng, submitted this Thesis for approval on 2017-04-14 at 16:27.","This Thesis was approved for publication on 2017-04-18 at 11:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10768 on 2017-08-10 at 15:05:37","Made available in DSpace on 2017-08-10T20:32:56Z (GMT). No. of bitstreams: 2 TENG-THESIS-2017.pdf: 486378 bytes, checksum: 7806fa81f8e97de0ae61fb5e33e81115 (MD5) LICENSE.txt: 4204 bytes, checksum: 8a65a459f7723bf11d5082e14db4b63a (MD5) Previous issue date: 2017-04-18","Embargo set by: Colleen Fallaw for item 102760 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102760 on 2019-08-11T09:15:24Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97707"],"dc:language":["en"],"dc:rights":["Copyright 2017 Bo Teng"],"dc:subject":["Straggler","Machine learning","Apache Spark","Iterative application"],"dc:title":["Mitigating Spark straggler tasks for iterative applications by data re-partitioning"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:34Z"}