{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113017"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113017","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Detecting, characterizing, and taming flaky tests","abstract":"As software evolves, developers typically perform regression testing to ensure that their code changes do not break existing functionalities. During regression testing, developers can waste time debugging their code changes because of spurious failures from flaky tests, which are tests that non-deterministically pass or fail on the same code. These spurious failures mislead developers about their code changes because the failures are often due to bugs that existed before the code changes. One prominent category of flaky tests is order-dependent (OD) flaky tests. Each OD test has at least one order in which the test passes and another order in which the test fails, and for every test order, the test either passes or fails in all runs of that test order. Another prominent category is async-wait (AW) flaky tests. Each AW test makes at least one asynchronous call and passes if the asynchronous call finishes on time but fails if the call finishes too early or too late. This dissertation tackles three main aspects of flaky tests. First, this dissertation presents novel techniques to detect flaky tests so that developers can preemptively prevent the problem of flaky tests from affecting their regression testing results. Second, this dissertation presents novel techniques to characterize flaky tests to help developers better understand their flaky tests and to help researchers invent new solutions to the flaky-test problem. Lastly, this dissertation presents novel techniques to tame the problem of flaky tests by accommodating the flakiness so that flaky tests do not mislead developers during regression testing. For detecting flaky tests, this dissertation presents (1) iDFlakies, a framework for detecting and partially classifying flaky tests, and IDoFT, an increasingly used dataset of flaky tests found in popular open-source projects; (2) an analysis of the probability to detect OD tests from randomizing test orders and a novel algorithm to systematically explore all consecutive pairs of tests, guaranteeing to detect all OD tests that depend on one other test; and (3) a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky—the results provide guidelines for when and how developers should spend their efforts to detect flaky tests. For characterizing flaky tests, this dissertation presents (1) the first automated technique to help developers debug flaky-test failures; and (2) a study to understand the effect that test orders have on non-deterministic tests, which can pass and fail even for the same test order—the results suggest that many of these tests can fail with significantly different failure rates for different test orders. Lastly, for taming flaky tests, this dissertation presents (1) the first automated techniques to reduce the number of spurious failures from OD tests, reducing such failures by 73%; and (2) the first automated techniques to speed up AW flaky tests while also keeping the number of spurious failures low, speeding up such tests by 38%. Overall, the work in this dissertation has helped detect more than 2000 flaky tests in over 150 open-source projects and fix more than 500 flaky tests in over 80 open-source projects.","abstract_html":"As software evolves, developers typically perform regression testing to ensure that their code changes do not break existing functionalities. During regression testing, developers can waste time debugging their code changes because of spurious failures from flaky tests, which are tests that non-deterministically pass or fail on the same code. These spurious failures mislead developers about their code changes because the failures are often due to bugs that existed before the code changes. One prominent category of flaky tests is order-dependent (OD) flaky tests. Each OD test has at least one order in which the test passes and another order in which the test fails, and for every test order, the test either passes or fails in all runs of that test order. Another prominent category is async-wait (AW) flaky tests. Each AW test makes at least one asynchronous call and passes if the asynchronous call finishes on time but fails if the call finishes too early or too late. This dissertation tackles three main aspects of flaky tests. First, this dissertation presents novel techniques to detect flaky tests so that developers can preemptively prevent the problem of flaky tests from affecting their regression testing results. Second, this dissertation presents novel techniques to characterize flaky tests to help developers better understand their flaky tests and to help researchers invent new solutions to the flaky-test problem. Lastly, this dissertation presents novel techniques to tame the problem of flaky tests by accommodating the flakiness so that flaky tests do not mislead developers during regression testing. For detecting flaky tests, this dissertation presents (1) iDFlakies, a framework for detecting and partially classifying flaky tests, and IDoFT, an increasingly used dataset of flaky tests found in popular open-source projects; (2) an analysis of the probability to detect OD tests from randomizing test orders and a novel algorithm to systematically explore all consecutive pairs of tests, guaranteeing to detect all OD tests that depend on one other test; and (3) a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky—the results provide guidelines for when and how developers should spend their efforts to detect flaky tests. For characterizing flaky tests, this dissertation presents (1) the first automated technique to help developers debug flaky-test failures; and (2) a study to understand the effect that test orders have on non-deterministic tests, which can pass and fail even for the same test order—the results suggest that many of these tests can fail with significantly different failure rates for different test orders. Lastly, for taming flaky tests, this dissertation presents (1) the first automated techniques to reduce the number of spurious failures from OD tests, reducing such failures by 73%; and (2) the first automated techniques to speed up AW flaky tests while also keeping the number of spurious failures low, speeding up such tests by 38%. Overall, the work in this dissertation has helped detect more than 2000 flaky tests in over 150 open-source projects and fix more than 500 flaky tests in over 80 open-source projects.","abstract_has_math":false,"creators":["Lam, Wing"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Xie, Tao","Marinov, Darko","Nath, Suman","Xu, Tianyin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:37Z","date_published":"2022-01-12T21:45:37Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Software engineering","software testing","regression testing","flaky tests"],"languages":["en"],"rights":["Copyright 2021 Wing Lam"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113017","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Xie, Tao","Marinov, Darko","Nath, Suman","Xu, Tianyin"]},{"key":"dc:creator","label":"Author","values":["Lam, Wing"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:37Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Software engineering","software testing","regression testing","flaky tests"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Wing Lam"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113017"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As software evolves, developers typically perform regression testing to ensure that their code changes do not break existing functionalities. During regression testing, developers can waste time debugging their code changes because of spurious failures from flaky tests, which are tests that non-deterministically pass or fail on the same code. These spurious failures mislead developers about their code changes because the failures are often due to bugs that existed before the code changes. One prominent category of flaky tests is order-dependent (OD) flaky tests. Each OD test has at least one order in which the test passes and another order in which the test fails, and for every test order, the test either passes or fails in all runs of that test order. Another prominent category is async-wait (AW) flaky tests. Each AW test makes at least one asynchronous call and passes if the asynchronous call finishes on time but fails if the call finishes too early or too late. This dissertation tackles three main aspects of flaky tests. First, this dissertation presents novel techniques to detect flaky tests so that developers can preemptively prevent the problem of flaky tests from affecting their regression testing results. Second, this dissertation presents novel techniques to characterize flaky tests to help developers better understand their flaky tests and to help researchers invent new solutions to the flaky-test problem. Lastly, this dissertation presents novel techniques to tame the problem of flaky tests by accommodating the flakiness so that flaky tests do not mislead developers during regression testing. For detecting flaky tests, this dissertation presents (1) iDFlakies, a framework for detecting and partially classifying flaky tests, and IDoFT, an increasingly used dataset of flaky tests found in popular open-source projects; (2) an analysis of the probability to detect OD tests from randomizing test orders and a novel algorithm to systematically explore all consecutive pairs of tests, guaranteeing to detect all OD tests that depend on one other test; and (3) a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky—the results provide guidelines for when and how developers should spend their efforts to detect flaky tests. For characterizing flaky tests, this dissertation presents (1) the first automated technique to help developers debug flaky-test failures; and (2) a study to understand the effect that test orders have on non-deterministic tests, which can pass and fail even for the same test order—the results suggest that many of these tests can fail with significantly different failure rates for different test orders. Lastly, for taming flaky tests, this dissertation presents (1) the first automated techniques to reduce the number of spurious failures from OD tests, reducing such failures by 73%; and (2) the first automated techniques to speed up AW flaky tests while also keeping the number of spurious failures low, speeding up such tests by 38%. Overall, the work in this dissertation has helped detect more than 2000 flaky tests in over 150 open-source projects and fix more than 500 flaky tests in over 80 open-source projects.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Wing Lam, accepted the attached license on 2021-07-12 at 00:07.","The student, Wing Lam, submitted this Dissertation for approval on 2021-07-12 at 00:18.","This Dissertation was approved for publication on 2021-07-12 at 14:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16854 on 2022-01-12 at 12:44:54","Made available in DSpace on 2022-01-12T21:45:37Z (GMT). No. of bitstreams: 3 LAM-DISSERTATION-2021.pdf: 2412160 bytes, checksum: ec45e5411611d304bf4b2f86163f0f7f (MD5) LICENSE.txt: 4205 bytes, checksum: 2510a50b6521d8c79124b9de0053d709 (MD5) PROQUEST_LICENSE.txt: 4551 bytes, checksum: 18c1a59a0d83295506399d57425fd193 (MD5) Previous issue date: 2021-07-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Detecting, characterizing, and taming flaky tests"]}]}],"canonical_facts":{"dc:contributor":["Xie, Tao","Marinov, Darko","Nath, Suman","Xu, Tianyin"],"dc:creator":["Lam, Wing"],"dc:date":["2022-01-12T21:45:37Z","2021-07-12","2021-08"],"dc:description":["As software evolves, developers typically perform regression testing to ensure that their code changes do not break existing functionalities. During regression testing, developers can waste time debugging their code changes because of spurious failures from flaky tests, which are tests that non-deterministically pass or fail on the same code. These spurious failures mislead developers about their code changes because the failures are often due to bugs that existed before the code changes. One prominent category of flaky tests is order-dependent (OD) flaky tests. Each OD test has at least one order in which the test passes and another order in which the test fails, and for every test order, the test either passes or fails in all runs of that test order. Another prominent category is async-wait (AW) flaky tests. Each AW test makes at least one asynchronous call and passes if the asynchronous call finishes on time but fails if the call finishes too early or too late. This dissertation tackles three main aspects of flaky tests. First, this dissertation presents novel techniques to detect flaky tests so that developers can preemptively prevent the problem of flaky tests from affecting their regression testing results. Second, this dissertation presents novel techniques to characterize flaky tests to help developers better understand their flaky tests and to help researchers invent new solutions to the flaky-test problem. Lastly, this dissertation presents novel techniques to tame the problem of flaky tests by accommodating the flakiness so that flaky tests do not mislead developers during regression testing. For detecting flaky tests, this dissertation presents (1) iDFlakies, a framework for detecting and partially classifying flaky tests, and IDoFT, an increasingly used dataset of flaky tests found in popular open-source projects; (2) an analysis of the probability to detect OD tests from randomizing test orders and a novel algorithm to systematically explore all consecutive pairs of tests, guaranteeing to detect all OD tests that depend on one other test; and (3) a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky—the results provide guidelines for when and how developers should spend their efforts to detect flaky tests. For characterizing flaky tests, this dissertation presents (1) the first automated technique to help developers debug flaky-test failures; and (2) a study to understand the effect that test orders have on non-deterministic tests, which can pass and fail even for the same test order—the results suggest that many of these tests can fail with significantly different failure rates for different test orders. Lastly, for taming flaky tests, this dissertation presents (1) the first automated techniques to reduce the number of spurious failures from OD tests, reducing such failures by 73%; and (2) the first automated techniques to speed up AW flaky tests while also keeping the number of spurious failures low, speeding up such tests by 38%. Overall, the work in this dissertation has helped detect more than 2000 flaky tests in over 150 open-source projects and fix more than 500 flaky tests in over 80 open-source projects.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Wing Lam, accepted the attached license on 2021-07-12 at 00:07.","The student, Wing Lam, submitted this Dissertation for approval on 2021-07-12 at 00:18.","This Dissertation was approved for publication on 2021-07-12 at 14:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16854 on 2022-01-12 at 12:44:54","Made available in DSpace on 2022-01-12T21:45:37Z (GMT). No. of bitstreams: 3 LAM-DISSERTATION-2021.pdf: 2412160 bytes, checksum: ec45e5411611d304bf4b2f86163f0f7f (MD5) LICENSE.txt: 4205 bytes, checksum: 2510a50b6521d8c79124b9de0053d709 (MD5) PROQUEST_LICENSE.txt: 4551 bytes, checksum: 18c1a59a0d83295506399d57425fd193 (MD5) Previous issue date: 2021-07-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113017"],"dc:language":["en"],"dc:rights":["Copyright 2021 Wing Lam"],"dc:subject":["Software engineering","software testing","regression testing","flaky tests"],"dc:title":["Detecting, characterizing, and taming flaky tests"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}