Back to results

University of Cambridge

Some Selective Inference and Optimization Methods for Reliable Causal Inference

Abstract

dc:description.abstract

In recent years, causal inference has seen growing uptake and research activity, attracted attention in new fields of applications and is deployed to analyse increasingly complex data. Adopting a causal perspective often allows us to get deeper insights into the data and the underlying system but typically relies on strong, sometimes overly optimistic assumptions. To ensure the validity of the obtained conclusions, it is thus crucial to develop methods to draw reliable inferences -- particularly in applications with high stakes. This thesis addresses this challenge in two contexts: adaptively collected experimental data and observational data with unmeasured confounding. It is organized in three chapters. The first chapter outlines the historical development of causal inference and reviews different frameworks for defining causality. We describe common approaches to draw inference from both experimental and observational data, discuss the required identification assumptions and situate our work in this context. The second chapter considers multi-stage adaptive experiments which use already collected data to determine the further course of action. Because the null hypothesis and experimental design are not pre-specified, it has long been recognized that statistical inference for adaptive experiments is not straightforward. Most existing methods only apply to specific adaptive designs and rely on strong assumptions. In this work, we propose 'selective randomization inference' as a general framework for analysing adaptive experiments. In a nutshell, our approach applies conditional post-selection inference to randomization tests. We propose a selective randomization p-value and prove that it both controls the selective type-I error and is computable, requiring neither independent or identically distributed data nor any other modelling assumptions. In order to compute a Monte Carlo approximation of the p-value, we employ rejection sampling or the random-walk Metropolis-Hastings algorithm. Furthermore, confidence intervals for a homogeneous treatment effect can be constructed via inversion of tests. To mitigate the risk of disconnected confidence intervals, we propose the use of hold-out units. Lastly, the chapter demonstrates the proposed method and compares it with other randomization tests using synthetic and real-world data. The third chapter is concerned with causal inference using observational data. In this setting, unmeasured confounding variables are a ubiquitous concern for drawing reliable conclusions. Hence, it is crucial to assess the robustness of the obtained results. However, many existing methods for such sensitivity analyses are limited to relatively simple models, difficult to interpret and consequently find only occasional use in practice. First, this chapter reviews the existing methodology and literature about sensitivity analysis and partial identification: Framing sensitivity analysis as a constrained stochastic optimization problem emerges as a flexible approach that allows us to specify interpretable sensitivity models. In this work, we leverage this perspective to conduct sensitivity analysis for a linear causal effect when an unmeasured confounder and a potential instrument are present. We show how the bias of the ordinary least squares (OLS) and two-stage least squares (TSLS) estimands can be expressed in terms of partial correlations. Leveraging the algebraic rules that relate different partial correlations, practitioners can specify intuitive sensitivity models which bound the bias. We further show that the heuristic ``plug-in'' sensitivity interval may not have any confidence guarantees; instead, we propose a bootstrap approach to construct sensitivity intervals. In order to visualize the results of the sensitivity analysis and different choices of the constraints on the sensitivity parameters, we employ contour plots. Lastly, the proposed methods are illustrated with a real study on the causal effect of education on earnings.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Freidling, Tobias
Advisor dc:contributor.advisor
  • Zhao, Qingyuan

Subjects

dc:subject × 3

Rights

dc:rights

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.118651
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/384748

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Freidling, Tobias. Some Selective Inference and Optimization Methods for Reliable Causal Inference. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.118651