{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110577"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110577","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"On the finite sample complexity of causal discovery and the value of domain expertise","abstract":"Causal discovery methods seek to identify causal relations between random variables from purely observational data, as opposed to actively collected experimental data where an experimenter intervenes on a subset of correlates. One of the seminal works in this area is the Inferred Causation (IC) algorithm, which guarantees successful causal discovery under the assumption of a conditional independence (CI) oracle: an oracle that can states whether two random variables are conditionally independent given another set of random variables. Practical implementations of this algorithm incorporate statistical tests for conditional independence, in place of a CI oracle. In this thesis, we analyze the sample complexity of causal discovery algorithms without a CI oracle: given a certain level of confidence, how many data points are needed for a causal discovery algorithm to identify a causal structure? Furthermore, our methods allow us to quantify the value of domain expertise in terms of data samples. Finally, we demonstrate the accuracy of these sample rates with numerical examples, and quantify the benefits of three types of domain expertise: sparsity priors, known causal directions, and known conditional dependencies.","abstract_html":"Causal discovery methods seek to identify causal relations between random variables from purely observational data, as opposed to actively collected experimental data where an experimenter intervenes on a subset of correlates. One of the seminal works in this area is the Inferred Causation (IC) algorithm, which guarantees successful causal discovery under the assumption of a conditional independence (CI) oracle: an oracle that can states whether two random variables are conditionally independent given another set of random variables. Practical implementations of this algorithm incorporate statistical tests for conditional independence, in place of a CI oracle. In this thesis, we analyze the sample complexity of causal discovery algorithms without a CI oracle: given a certain level of confidence, how many data points are needed for a causal discovery algorithm to identify a causal structure? Furthermore, our methods allow us to quantify the value of domain expertise in terms of data samples. Finally, we demonstrate the accuracy of these sample rates with numerical examples, and quantify the benefits of three types of domain expertise: sparsity priors, known causal directions, and known conditional dependencies.","abstract_has_math":false,"creators":["Wadhwa, Samir"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Mechanical Engineering","degree_department":null,"school":null,"contributors":["Dong, Roy","Dullerud, Geir Eirik"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T01:13:28Z","date_published":"2021-09-17T01:13:28Z","updated_at":"2026-07-22T22:24:52Z","subjects":["causality","discrete Bayesian networks","conditional independence testing","family-wise error rate"],"languages":["en"],"rights":["Copyright 2021 Samir Wadhwa"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110577","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dong, Roy","Dullerud, Geir Eirik"]},{"key":"dc:creator","label":"Author","values":["Wadhwa, Samir"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T01:13:28Z","2021-04-28","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechanical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["causality","discrete Bayesian networks","conditional independence testing","family-wise error rate"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Samir Wadhwa"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110577"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Causal discovery methods seek to identify causal relations between random variables from purely observational data, as opposed to actively collected experimental data where an experimenter intervenes on a subset of correlates. One of the seminal works in this area is the Inferred Causation (IC) algorithm, which guarantees successful causal discovery under the assumption of a conditional independence (CI) oracle: an oracle that can states whether two random variables are conditionally independent given another set of random variables. Practical implementations of this algorithm incorporate statistical tests for conditional independence, in place of a CI oracle. In this thesis, we analyze the sample complexity of causal discovery algorithms without a CI oracle: given a certain level of confidence, how many data points are needed for a causal discovery algorithm to identify a causal structure? Furthermore, our methods allow us to quantify the value of domain expertise in terms of data samples. Finally, we demonstrate the accuracy of these sample rates with numerical examples, and quantify the benefits of three types of domain expertise: sparsity priors, known causal directions, and known conditional dependencies.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Samir Wadhwa, accepted the attached license on 2021-04-26 at 13:32.","The student, Samir Wadhwa, submitted this Thesis for approval on 2021-04-26 at 13:44.","This Thesis was approved for publication on 2021-04-28 at 14:20.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16561 on 2021-09-16 at 16:48:12","Made available in DSpace on 2021-09-17T01:13:28Z (GMT). No. of bitstreams: 2 WADHWA-THESIS-2021.pdf: 560728 bytes, checksum: d36b252b3361a6dbb523b91fbbe946f9 (MD5) LICENSE.txt: 4209 bytes, checksum: d65985c7f828d5b980ebc41aa7c46609 (MD5) Previous issue date: 2021-04-28"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["On the finite sample complexity of causal discovery and the value of domain expertise"]}]}],"canonical_facts":{"dc:contributor":["Dong, Roy","Dullerud, Geir Eirik"],"dc:creator":["Wadhwa, Samir"],"dc:date":["2021-09-17T01:13:28Z","2021-04-28","2021-05"],"dc:description":["Causal discovery methods seek to identify causal relations between random variables from purely observational data, as opposed to actively collected experimental data where an experimenter intervenes on a subset of correlates. One of the seminal works in this area is the Inferred Causation (IC) algorithm, which guarantees successful causal discovery under the assumption of a conditional independence (CI) oracle: an oracle that can states whether two random variables are conditionally independent given another set of random variables. Practical implementations of this algorithm incorporate statistical tests for conditional independence, in place of a CI oracle. In this thesis, we analyze the sample complexity of causal discovery algorithms without a CI oracle: given a certain level of confidence, how many data points are needed for a causal discovery algorithm to identify a causal structure? Furthermore, our methods allow us to quantify the value of domain expertise in terms of data samples. Finally, we demonstrate the accuracy of these sample rates with numerical examples, and quantify the benefits of three types of domain expertise: sparsity priors, known causal directions, and known conditional dependencies.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Samir Wadhwa, accepted the attached license on 2021-04-26 at 13:32.","The student, Samir Wadhwa, submitted this Thesis for approval on 2021-04-26 at 13:44.","This Thesis was approved for publication on 2021-04-28 at 14:20.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16561 on 2021-09-16 at 16:48:12","Made available in DSpace on 2021-09-17T01:13:28Z (GMT). No. of bitstreams: 2 WADHWA-THESIS-2021.pdf: 560728 bytes, checksum: d36b252b3361a6dbb523b91fbbe946f9 (MD5) LICENSE.txt: 4209 bytes, checksum: d65985c7f828d5b980ebc41aa7c46609 (MD5) Previous issue date: 2021-04-28"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110577"],"dc:language":["en"],"dc:rights":["Copyright 2021 Samir Wadhwa"],"dc:subject":["causality","discrete Bayesian networks","conditional independence testing","family-wise error rate"],"dc:title":["On the finite sample complexity of causal discovery and the value of domain expertise"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Mechanical Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}