{"id":{"repo_id":"penn","oai_identifier":"oai:repository.upenn.edu:20.500.14332/60083"},"canonical_url":"https://search.dev.ndltd.org/etd/penn/oai:repository.upenn.edu:20.500.14332/60083","repository":{"repo_id":"penn","name":"University of Pennsylvania","base_url":"https://repository.upenn.edu/server/oai/request"},"display":{"title":"Improving Observational Causality Using Machine Learning","abstract":"Causality is at the heart of many machine learning questions whether we know it or not, and we need to explicitly incorporate causal reasoning in order to answer them effectively. By a similar token, traditional causal inference methods can benefit from machine learning to adapt to more complex data domains. This thesis will explore the interplay between observational causal inference and machine learning, focusing on improving different aspects of the causal inference study lifecycle. Namely, we develop methods that facilitate the discovery of new study opportunities, improve the feasibility of existing studies, and allow for better interpretation of the resulting causal estimates. As identifying causal inference opportunities is currently a manual process requiring human intuition, we first develop a scaleable method for data-driven discovery of regression discontinuities, a class of observational causal inference methods. Next, we re-frame observational study exclusion criteria as a well-posed machine learning task, increasing interpretability by characterizing the excluded units. Both our discovery and exclusion criteria methods explicitly account for maximizing statistical power to increase study feasibility, and both are evaluated for their real-world efficacy on a medical claims dataset with over 60 million patients. Finally, we show the utility of incorporating machine learning into the causal study lifecycle through a large-scale study of the impact of civility in online social interactions. Through these works, we highlight not only how machine learning can improve causal inference in observational data settings but also the need to consider causality across traditional machine learning tasks.","abstract_html":"Causality is at the heart of many machine learning questions whether we know it or not, and we need to explicitly incorporate causal reasoning in order to answer them effectively. By a similar token, traditional causal inference methods can benefit from machine learning to adapt to more complex data domains. This thesis will explore the interplay between observational causal inference and machine learning, focusing on improving different aspects of the causal inference study lifecycle. Namely, we develop methods that facilitate the discovery of new study opportunities, improve the feasibility of existing studies, and allow for better interpretation of the resulting causal estimates. As identifying causal inference opportunities is currently a manual process requiring human intuition, we first develop a scaleable method for data-driven discovery of regression discontinuities, a class of observational causal inference methods. Next, we re-frame observational study exclusion criteria as a well-posed machine learning task, increasing interpretability by characterizing the excluded units. Both our discovery and exclusion criteria methods explicitly account for maximizing statistical power to increase study feasibility, and both are evaluated for their real-world efficacy on a medical claims dataset with over 60 million patients. Finally, we show the utility of incorporating machine learning into the causal study lifecycle through a large-scale study of the impact of civility in online social interactions. Through these works, we highlight not only how machine learning can improve causal inference in observational data settings but also the need to consider causality across traditional machine learning tasks.","abstract_has_math":false,"creators":["Liu, Tong"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Ungar, Lyle, H","Kording, Konrad, P"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-24T03:47:40Z","subjects":["Computer Sciences","Public Health","Statistics and Probability"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://repository.upenn.edu/handle/20.500.14332/60083","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Ungar, Lyle, H","Kording, Konrad, P"]},{"key":"dc:creator","label":"Author","values":["Liu, Tong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-06-18T14:18:38Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-06-18T14:18:38Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation/Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Sciences","Public Health","Statistics and Probability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://repository.upenn.edu/handle/20.500.14332/60083"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Causality is at the heart of many machine learning questions whether we know it or not, and we need to explicitly incorporate causal reasoning in order to answer them effectively. By a similar token, traditional causal inference methods can benefit from machine learning to adapt to more complex data domains. This thesis will explore the interplay between observational causal inference and machine learning, focusing on improving different aspects of the causal inference study lifecycle. Namely, we develop methods that facilitate the discovery of new study opportunities, improve the feasibility of existing studies, and allow for better interpretation of the resulting causal estimates. As identifying causal inference opportunities is currently a manual process requiring human intuition, we first develop a scaleable method for data-driven discovery of regression discontinuities, a class of observational causal inference methods. Next, we re-frame observational study exclusion criteria as a well-posed machine learning task, increasing interpretability by characterizing the excluded units. Both our discovery and exclusion criteria methods explicitly account for maximizing statistical power to increase study feasibility, and both are evaluated for their real-world efficacy on a medical claims dataset with over 60 million patients. Finally, we show the utility of incorporating machine learning into the causal study lifecycle through a large-scale study of the impact of civility in online social interactions. Through these works, we highlight not only how machine learning can improve causal inference in observational data settings but also the need to consider causality across traditional machine learning tasks."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy (PhD)"]},{"key":"dc:title","label":"Title","values":["Improving Observational Causality Using Machine Learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Ungar, Lyle, H","Kording, Konrad, P"],"dc:creator":["Liu, Tong"],"dc:date.accessioned":["2024-06-18T14:18:38Z"],"dc:date.available":["2024-06-18T14:18:38Z"],"dc:date.issued":["2024"],"dc:description.abstract":["Causality is at the heart of many machine learning questions whether we know it or not, and we need to explicitly incorporate causal reasoning in order to answer them effectively. By a similar token, traditional causal inference methods can benefit from machine learning to adapt to more complex data domains. This thesis will explore the interplay between observational causal inference and machine learning, focusing on improving different aspects of the causal inference study lifecycle. Namely, we develop methods that facilitate the discovery of new study opportunities, improve the feasibility of existing studies, and allow for better interpretation of the resulting causal estimates. As identifying causal inference opportunities is currently a manual process requiring human intuition, we first develop a scaleable method for data-driven discovery of regression discontinuities, a class of observational causal inference methods. Next, we re-frame observational study exclusion criteria as a well-posed machine learning task, increasing interpretability by characterizing the excluded units. Both our discovery and exclusion criteria methods explicitly account for maximizing statistical power to increase study feasibility, and both are evaluated for their real-world efficacy on a medical claims dataset with over 60 million patients. Finally, we show the utility of incorporating machine learning into the causal study lifecycle through a large-scale study of the impact of civility in online social interactions. Through these works, we highlight not only how machine learning can improve causal inference in observational data settings but also the need to consider causality across traditional machine learning tasks."],"dc:description.degree":["Doctor of Philosophy (PhD)"],"dc:identifier.uri":["https://repository.upenn.edu/handle/20.500.14332/60083"],"dc:language.iso":["en"],"dc:subject":["Computer Sciences","Public Health","Statistics and Probability"],"dc:title":["Improving Observational Causality Using Machine Learning"],"dc:type":["Dissertation/Thesis"]},"updated_at":"2026-07-24T03:47:40Z"}