University of Cambridge
Robust and interpretable high-dimensional machine learning for predictive cancer medicine
Abstract
dc:description.abstractGaining actionable insights from complex, high-dimensional biological data is challenging and often relies upon dimensionality reduction techniques. These methods reveal structure within data, potentially exposing meaningful biological patterns. In predictive cancer medicine, it is common to employ linear dimensionality reduction methods due to their inherent interpretability. However, recent advances in machine learning have sparked considerable interest in more flexible, non-linear dimensionality reduction techniques. In this thesis, I propose and build upon extensions to the variational autoencoder (VAE), a probabilistic latent variable model that leverages neural networks to learn latent representations. Specifically, I propose a VAE variant that generates latent representations which explicitly capture genetic dependencies in cancers. I incorporate this into a framework for prediction using these learned interpretable representations as inputs. I demonstrate that this process allows interpretable and biologically meaningful prediction on a variety of tasks in cancer medicine. Extending this two-stage learning framework, I propose a series of models which instead jointly learn interpretable representations and downstream prediction models while effectively leveraging informative auxiliary data to encourage the formation of useful representations. Complementing these latent variable modelling approaches, I also explore interpretability directly within the original high-dimensional feature space using interpretable graph neural networks (GNNs). In particular, I focus on developing interpretable GNNs that are suited to problems involving large biomedical graphs. I apply these methods to identify genetic dependencies and biological interactions that drive drug response in cancer cell lines. These GNN-based methods provide transparent explanations by explicitly selecting and highlighting relevant features from high-dimensional data projected onto large biomedical graphs. I demonstrate the effectiveness and interpretability of the proposed methodologies through extensive evaluation on diverse datasets, including high-dimensional transcriptomics data from cancer cell lines, high-throughput in vitro screening datasets, and clinical cancer cohorts, as well as simulated datasets. These approaches not only improve predictive performance and generalisation but also provide transparent, biologically interpretable insights, demonstrating their potential utility in personalised cancer medicine.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Kirkham, Dominic
- Advisors dc:contributor.advisor
-
- Mukherjee, Sach
- Rueda Palacio, Oscar
Subjects
dc:subject × 2Rights
dc:rights- Licence
- Language dc:language
- eng
Identifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.124927
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/395408