{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/389961"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/389961","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Deep representation learning for cancer transcriptomics: towards robust models for heterogeneous data","abstract":"Modern machine learning (ML) methods have the potential to transform research and clinical practice across biomedicine, including cancer diagnosis and treatment. Molecular data, including genomic and transcriptomic profiles, are used to train these ML systems. But current models are very sensitive to distributional shifts in the underlying training data, which arise from variations in collection, preservation and processing methods. This severely limits model generalisability across clinical settings, making it difficult to deploy them at scale in national healthcare systems. This thesis investigates three deep learning approaches to address these challenges for transcriptomic data. First, we develop DeepRENEW, a self-supervised autoencoder that aims to mitigate the effects of noise introduced by different data-acquisition techniques. Second, we explore architectural extensions to DeepRENEW inspired by neural style transfer. We consider a number of different methods to explicitly model transitions between preservation modalities, offering valuable methodological insights into cancer transcriptomic transfer learning, despite not finding a method that consistently improves the original model. Third, we introduce a novel semi-supervised causal regularisation approach for cancer survival prediction. This integrates observational patient data with interventional CRISPR-interference experiments to enhance survival prediction robustness. We find that deep learning holds great promise for addressing distributional shifts in cancer transcriptomics, but that much larger, more diverse datasets are needed, especially to train future biological foundation models for bulk transcriptome modelling.","abstract_html":"Modern machine learning (ML) methods have the potential to transform research and clinical practice across biomedicine, including cancer diagnosis and treatment. Molecular data, including genomic and transcriptomic profiles, are used to train these ML systems. But current models are very sensitive to distributional shifts in the underlying training data, which arise from variations in collection, preservation and processing methods. This severely limits model generalisability across clinical settings, making it difficult to deploy them at scale in national healthcare systems. This thesis investigates three deep learning approaches to address these challenges for transcriptomic data. First, we develop DeepRENEW, a self-supervised autoencoder that aims to mitigate the effects of noise introduced by different data-acquisition techniques. Second, we explore architectural extensions to DeepRENEW inspired by neural style transfer. We consider a number of different methods to explicitly model transitions between preservation modalities, offering valuable methodological insights into cancer transcriptomic transfer learning, despite not finding a method that consistently improves the original model. Third, we introduce a novel semi-supervised causal regularisation approach for cancer survival prediction. This integrates observational patient data with interventional CRISPR-interference experiments to enhance survival prediction robustness. We find that deep learning holds great promise for addressing distributional shifts in cancer transcriptomics, but that much larger, more diverse datasets are needed, especially to train future biological foundation models for bulk transcriptome modelling.","abstract_has_math":false,"creators":["Moulange, Richard"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Mukherjee, Sach","Rueda, Oscar M"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-31","date_published":"2025-03-31","updated_at":"2026-07-22T22:23:59Z","subjects":["Deep learning","Cancer","Transcriptomics"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/343a8bb9-a531-41de-b873-bd1cd4f9a929/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.121686","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Mukherjee, Sach","Rueda, Oscar M"]},{"key":"dc:creator","label":"Author","values":["Moulange, Richard"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-03-31"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/389961"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Deep learning","Cancer","Transcriptomics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/343a8bb9-a531-41de-b873-bd1cd4f9a929/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.121686"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/eea64883-04b4-4170-a327-89cda1a6086f/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Modern machine learning (ML) methods have the potential to transform research and clinical practice across biomedicine, including cancer diagnosis and treatment. Molecular data, including genomic and transcriptomic profiles, are used to train these ML systems. But current models are very sensitive to distributional shifts in the underlying training data, which arise from variations in collection, preservation and processing methods. This severely limits model generalisability across clinical settings, making it difficult to deploy them at scale in national healthcare systems. This thesis investigates three deep learning approaches to address these challenges for transcriptomic data. First, we develop DeepRENEW, a self-supervised autoencoder that aims to mitigate the effects of noise introduced by different data-acquisition techniques. Second, we explore architectural extensions to DeepRENEW inspired by neural style transfer. We consider a number of different methods to explicitly model transitions between preservation modalities, offering valuable methodological insights into cancer transcriptomic transfer learning, despite not finding a method that consistently improves the original model. Third, we introduce a novel semi-supervised causal regularisation approach for cancer survival prediction. This integrates observational patient data with interventional CRISPR-interference experiments to enhance survival prediction robustness. We find that deep learning holds great promise for addressing distributional shifts in cancer transcriptomics, but that much larger, more diverse datasets are needed, especially to train future biological foundation models for bulk transcriptome modelling."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["ef4d383b7d2ec5f1bc69670443248a67","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Deep representation learning for cancer transcriptomics: towards robust models for heterogeneous data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Mukherjee, Sach","Rueda, Oscar M"],"dc:creator":["Moulange, Richard"],"dc:date.issued":["2025-03-31"],"dc:description.abstract":["Modern machine learning (ML) methods have the potential to transform research and clinical practice across biomedicine, including cancer diagnosis and treatment. Molecular data, including genomic and transcriptomic profiles, are used to train these ML systems. But current models are very sensitive to distributional shifts in the underlying training data, which arise from variations in collection, preservation and processing methods. This severely limits model generalisability across clinical settings, making it difficult to deploy them at scale in national healthcare systems. This thesis investigates three deep learning approaches to address these challenges for transcriptomic data. First, we develop DeepRENEW, a self-supervised autoencoder that aims to mitigate the effects of noise introduced by different data-acquisition techniques. Second, we explore architectural extensions to DeepRENEW inspired by neural style transfer. We consider a number of different methods to explicitly model transitions between preservation modalities, offering valuable methodological insights into cancer transcriptomic transfer learning, despite not finding a method that consistently improves the original model. Third, we introduce a novel semi-supervised causal regularisation approach for cancer survival prediction. This integrates observational patient data with interventional CRISPR-interference experiments to enhance survival prediction robustness. We find that deep learning holds great promise for addressing distributional shifts in cancer transcriptomics, but that much larger, more diverse datasets are needed, especially to train future biological foundation models for bulk transcriptome modelling."],"dc:format.checksum.md5":["ef4d383b7d2ec5f1bc69670443248a67","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.121686"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/eea64883-04b4-4170-a327-89cda1a6086f/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/389961"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/343a8bb9-a531-41de-b873-bd1cd4f9a929/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"dc:subject":["Deep learning","Cancer","Transcriptomics"],"dc:title":["Deep representation learning for cancer transcriptomics: towards robust models for heterogeneous data"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:59Z"}