{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/395408"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/395408","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Robust and interpretable high-dimensional machine learning for predictive cancer medicine","abstract":"Gaining actionable insights from complex, high-dimensional biological data is challenging and often relies upon dimensionality reduction techniques. These methods reveal structure within data, potentially exposing meaningful biological patterns. In predictive cancer medicine, it is common to employ linear dimensionality reduction methods due to their inherent interpretability. However, recent advances in machine learning have sparked considerable interest in more flexible, non-linear dimensionality reduction techniques. In this thesis, I propose and build upon extensions to the variational autoencoder (VAE), a probabilistic latent variable model that leverages neural networks to learn latent representations. Specifically, I propose a VAE variant that generates latent representations which explicitly capture genetic dependencies in cancers. I incorporate this into a framework for prediction using these learned interpretable representations as inputs. I demonstrate that this process allows interpretable and biologically meaningful prediction on a variety of tasks in cancer medicine. Extending this two-stage learning framework, I propose a series of models which instead jointly learn interpretable representations and downstream prediction models while effectively leveraging informative auxiliary data to encourage the formation of useful representations. Complementing these latent variable modelling approaches, I also explore interpretability directly within the original high-dimensional feature space using interpretable graph neural networks (GNNs). In particular, I focus on developing interpretable GNNs that are suited to problems involving large biomedical graphs. I apply these methods to identify genetic dependencies and biological interactions that drive drug response in cancer cell lines. These GNN-based methods provide transparent explanations by explicitly selecting and highlighting relevant features from high-dimensional data projected onto large biomedical graphs. I demonstrate the effectiveness and interpretability of the proposed methodologies through extensive evaluation on diverse datasets, including high-dimensional transcriptomics data from cancer cell lines, high-throughput in vitro screening datasets, and clinical cancer cohorts, as well as simulated datasets. These approaches not only improve predictive performance and generalisation but also provide transparent, biologically interpretable insights, demonstrating their potential utility in personalised cancer medicine.","abstract_html":"Gaining actionable insights from complex, high-dimensional biological data is challenging and often relies upon dimensionality reduction techniques. These methods reveal structure within data, potentially exposing meaningful biological patterns. In predictive cancer medicine, it is common to employ linear dimensionality reduction methods due to their inherent interpretability. However, recent advances in machine learning have sparked considerable interest in more flexible, non-linear dimensionality reduction techniques. In this thesis, I propose and build upon extensions to the variational autoencoder (VAE), a probabilistic latent variable model that leverages neural networks to learn latent representations. Specifically, I propose a VAE variant that generates latent representations which explicitly capture genetic dependencies in cancers. I incorporate this into a framework for prediction using these learned interpretable representations as inputs. I demonstrate that this process allows interpretable and biologically meaningful prediction on a variety of tasks in cancer medicine. Extending this two-stage learning framework, I propose a series of models which instead jointly learn interpretable representations and downstream prediction models while effectively leveraging informative auxiliary data to encourage the formation of useful representations. Complementing these latent variable modelling approaches, I also explore interpretability directly within the original high-dimensional feature space using interpretable graph neural networks (GNNs). In particular, I focus on developing interpretable GNNs that are suited to problems involving large biomedical graphs. I apply these methods to identify genetic dependencies and biological interactions that drive drug response in cancer cell lines. These GNN-based methods provide transparent explanations by explicitly selecting and highlighting relevant features from high-dimensional data projected onto large biomedical graphs. I demonstrate the effectiveness and interpretability of the proposed methodologies through extensive evaluation on diverse datasets, including high-dimensional transcriptomics data from cancer cell lines, high-throughput in vitro screening datasets, and clinical cancer cohorts, as well as simulated datasets. These approaches not only improve predictive performance and generalisation but also provide transparent, biologically interpretable insights, demonstrating their potential utility in personalised cancer medicine.","abstract_has_math":false,"creators":["Kirkham, Dominic"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Mukherjee, Sach","Rueda Palacio, Oscar"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-27","date_published":"2025-07-27","updated_at":"2026-07-22T22:24:01Z","subjects":["Cancer medicine","Machine learning"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/008b1ce1-d5f3-4bc8-b83d-2f8f71539dc1/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.124927","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Mukherjee, Sach","Rueda Palacio, Oscar"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Medical Research Council"]},{"key":"dc:creator","label":"Author","values":["Kirkham, Dominic"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-07-27"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/395408"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cancer medicine","Machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/008b1ce1-d5f3-4bc8-b83d-2f8f71539dc1/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.124927"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/32f6ee0b-879e-41bd-bd00-11902e14be19/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Gaining actionable insights from complex, high-dimensional biological data is challenging and often relies upon dimensionality reduction techniques. These methods reveal structure within data, potentially exposing meaningful biological patterns. In predictive cancer medicine, it is common to employ linear dimensionality reduction methods due to their inherent interpretability. However, recent advances in machine learning have sparked considerable interest in more flexible, non-linear dimensionality reduction techniques. In this thesis, I propose and build upon extensions to the variational autoencoder (VAE), a probabilistic latent variable model that leverages neural networks to learn latent representations. Specifically, I propose a VAE variant that generates latent representations which explicitly capture genetic dependencies in cancers. I incorporate this into a framework for prediction using these learned interpretable representations as inputs. I demonstrate that this process allows interpretable and biologically meaningful prediction on a variety of tasks in cancer medicine. Extending this two-stage learning framework, I propose a series of models which instead jointly learn interpretable representations and downstream prediction models while effectively leveraging informative auxiliary data to encourage the formation of useful representations. Complementing these latent variable modelling approaches, I also explore interpretability directly within the original high-dimensional feature space using interpretable graph neural networks (GNNs). In particular, I focus on developing interpretable GNNs that are suited to problems involving large biomedical graphs. I apply these methods to identify genetic dependencies and biological interactions that drive drug response in cancer cell lines. These GNN-based methods provide transparent explanations by explicitly selecting and highlighting relevant features from high-dimensional data projected onto large biomedical graphs. I demonstrate the effectiveness and interpretability of the proposed methodologies through extensive evaluation on diverse datasets, including high-dimensional transcriptomics data from cancer cell lines, high-throughput in vitro screening datasets, and clinical cancer cohorts, as well as simulated datasets. These approaches not only improve predictive performance and generalisation but also provide transparent, biologically interpretable insights, demonstrating their potential utility in personalised cancer medicine."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["ef495388e0e685257a9dd33940fb583f","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Robust and interpretable high-dimensional machine learning for predictive cancer medicine"]}]}],"canonical_facts":{"dc:contributor.advisor":["Mukherjee, Sach","Rueda Palacio, Oscar"],"dc:contributor.sponsor":["Medical Research Council"],"dc:creator":["Kirkham, Dominic"],"dc:date.issued":["2025-07-27"],"dc:description.abstract":["Gaining actionable insights from complex, high-dimensional biological data is challenging and often relies upon dimensionality reduction techniques. These methods reveal structure within data, potentially exposing meaningful biological patterns. In predictive cancer medicine, it is common to employ linear dimensionality reduction methods due to their inherent interpretability. However, recent advances in machine learning have sparked considerable interest in more flexible, non-linear dimensionality reduction techniques. In this thesis, I propose and build upon extensions to the variational autoencoder (VAE), a probabilistic latent variable model that leverages neural networks to learn latent representations. Specifically, I propose a VAE variant that generates latent representations which explicitly capture genetic dependencies in cancers. I incorporate this into a framework for prediction using these learned interpretable representations as inputs. I demonstrate that this process allows interpretable and biologically meaningful prediction on a variety of tasks in cancer medicine. Extending this two-stage learning framework, I propose a series of models which instead jointly learn interpretable representations and downstream prediction models while effectively leveraging informative auxiliary data to encourage the formation of useful representations. Complementing these latent variable modelling approaches, I also explore interpretability directly within the original high-dimensional feature space using interpretable graph neural networks (GNNs). In particular, I focus on developing interpretable GNNs that are suited to problems involving large biomedical graphs. I apply these methods to identify genetic dependencies and biological interactions that drive drug response in cancer cell lines. These GNN-based methods provide transparent explanations by explicitly selecting and highlighting relevant features from high-dimensional data projected onto large biomedical graphs. I demonstrate the effectiveness and interpretability of the proposed methodologies through extensive evaluation on diverse datasets, including high-dimensional transcriptomics data from cancer cell lines, high-throughput in vitro screening datasets, and clinical cancer cohorts, as well as simulated datasets. These approaches not only improve predictive performance and generalisation but also provide transparent, biologically interpretable insights, demonstrating their potential utility in personalised cancer medicine."],"dc:format.checksum.md5":["ef495388e0e685257a9dd33940fb583f","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.124927"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/32f6ee0b-879e-41bd-bd00-11902e14be19/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/395408"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/008b1ce1-d5f3-4bc8-b83d-2f8f71539dc1/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Cancer medicine","Machine learning"],"dc:title":["Robust and interpretable high-dimensional machine learning for predictive cancer medicine"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:01Z"}