{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/267784"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/267784","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Latent feature models and non-invasive clonal reconstruction","abstract":"Intratumoural heterogeneity complicates the molecular interpretation of biopsies, as multiple distinct tumour genomes are sampled and analysed at once. Ignoring the presence of these populations can lead to erroneous conclusions, and so a correct analysis must account for the clonal structure of the sample. Several methods to reconstruct tumour clonality from sequencing data have been proposed, spanning methods that either do not consider phylogenetic constraints or posit a perfect phylogeny. Models of the first type are typically latent feature models that can describe the observed data flexibly, but whose results may not be reconcilable with a phylogeny. The second type, instead, generally comprises non-parametric mixture models, with strict assumptions on the tumour’s evolutionary process. The focus of this dissertation is on the development of a phylogenetic latent feature model that can bridge the advantages of these two approaches, allowing deviations from a perfect phylogeny. The work is recounted by three statistical models of increasing complexity. First, I present a non-parametric model based on the Indian Buffet Process prior, and highlight the need for phylogenetic constraints. Second, I develop a finite, phylogenetic extension of the previous model, and show that it can outperform competing methods. Third, I generalise the phylogenetic model to arbitrary copy-number states. Markov chain Monte Carlo algorithms are presented to perform inference. The models are tested on datasets that include synthetic data, controlled biological data, and clinical data. In particular, the copy-number generalisation is applied to longitudinal circulating tumour DNA samples. Liquid biopsies that leverage circulating tumour DNA require sensitive techniques in order to detect mutations at low allele fractions. One method that allows sensitive mutation calling is the amplicon sequencing strategy TAm-Seq. I present bioinformatic tools to improve both the development of TAm-Seq amplicon panels and the analysis of its sequencing data. Finally, an enhancement of this method is presented and shown to detect mutations de novo and in a multiplexed manner at allele fractions less than 0.1%.","abstract_html":"Intratumoural heterogeneity complicates the molecular interpretation of biopsies, as multiple distinct tumour genomes are sampled and analysed at once. Ignoring the presence of these populations can lead to erroneous conclusions, and so a correct analysis must account for the clonal structure of the sample. Several methods to reconstruct tumour clonality from sequencing data have been proposed, spanning methods that either do not consider phylogenetic constraints or posit a perfect phylogeny. Models of the first type are typically latent feature models that can describe the observed data flexibly, but whose results may not be reconcilable with a phylogeny. The second type, instead, generally comprises non-parametric mixture models, with strict assumptions on the tumour’s evolutionary process. The focus of this dissertation is on the development of a phylogenetic latent feature model that can bridge the advantages of these two approaches, allowing deviations from a perfect phylogeny. The work is recounted by three statistical models of increasing complexity. First, I present a non-parametric model based on the Indian Buffet Process prior, and highlight the need for phylogenetic constraints. Second, I develop a finite, phylogenetic extension of the previous model, and show that it can outperform competing methods. Third, I generalise the phylogenetic model to arbitrary copy-number states. Markov chain Monte Carlo algorithms are presented to perform inference. The models are tested on datasets that include synthetic data, controlled biological data, and clinical data. In particular, the copy-number generalisation is applied to longitudinal circulating tumour DNA samples. Liquid biopsies that leverage circulating tumour DNA require sensitive techniques in order to detect mutations at low allele fractions. One method that allows sensitive mutation calling is the amplicon sequencing strategy TAm-Seq. I present bioinformatic tools to improve both the development of TAm-Seq amplicon panels and the analysis of its sequencing data. Finally, an enhancement of this method is presented and shown to detect mutations de novo and in a multiplexed manner at allele fractions less than 0.1%.","abstract_has_math":false,"creators":["Marass, Francesco"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Rosenfeld, Nitzan"],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-07-22","date_published":"2017-07-22","updated_at":"2026-07-22T22:24:03Z","subjects":["Cancer genomics","Tumour heterogeneity","Computational biology","Statistical modelling","Latent feature models","Deconvolution","Amplicon sequencing"],"languages":["en"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/17fae8f7-130d-44f5-8cd2-0e7f1bd33ac2/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000289937320"],"render_values":[{"text":"0000-0002-8993-7320","href":"https://orcid.org/0000-0002-8993-7320","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.13715","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Rosenfeld, Nitzan"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Cancer Research UK"]},{"key":"dc:creator","label":"Author","values":["Marass, Francesco"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000289937320"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2017-07-22"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/267784"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cancer genomics","Tumour heterogeneity","Computational biology","Statistical modelling","Latent feature models","Deconvolution","Amplicon sequencing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/17fae8f7-130d-44f5-8cd2-0e7f1bd33ac2/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.17863/CAM.13715"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/dedc20c6-cc5e-466d-a7bd-61fa962c7749/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Intratumoural heterogeneity complicates the molecular interpretation of biopsies, as multiple distinct tumour genomes are sampled and analysed at once. Ignoring the presence of these populations can lead to erroneous conclusions, and so a correct analysis must account for the clonal structure of the sample. Several methods to reconstruct tumour clonality from sequencing data have been proposed, spanning methods that either do not consider phylogenetic constraints or posit a perfect phylogeny. Models of the first type are typically latent feature models that can describe the observed data flexibly, but whose results may not be reconcilable with a phylogeny. The second type, instead, generally comprises non-parametric mixture models, with strict assumptions on the tumour’s evolutionary process. The focus of this dissertation is on the development of a phylogenetic latent feature model that can bridge the advantages of these two approaches, allowing deviations from a perfect phylogeny. The work is recounted by three statistical models of increasing complexity. First, I present a non-parametric model based on the Indian Buffet Process prior, and highlight the need for phylogenetic constraints. Second, I develop a finite, phylogenetic extension of the previous model, and show that it can outperform competing methods. Third, I generalise the phylogenetic model to arbitrary copy-number states. Markov chain Monte Carlo algorithms are presented to perform inference. The models are tested on datasets that include synthetic data, controlled biological data, and clinical data. In particular, the copy-number generalisation is applied to longitudinal circulating tumour DNA samples. Liquid biopsies that leverage circulating tumour DNA require sensitive techniques in order to detect mutations at low allele fractions. One method that allows sensitive mutation calling is the amplicon sequencing strategy TAm-Seq. I present bioinformatic tools to improve both the development of TAm-Seq amplicon panels and the analysis of its sequencing data. Finally, an enhancement of this method is presented and shown to detect mutations de novo and in a multiplexed manner at allele fractions less than 0.1%."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["87eda9de84448d1f82354d60eee3eb5f","c4e826f60430ccf66e2129c9c012e52f"]},{"key":"dc:title","label":"Title","values":["Latent feature models and non-invasive clonal reconstruction"]}]}],"canonical_facts":{"dc:contributor.advisor":["Rosenfeld, Nitzan"],"dc:contributor.sponsor":["Cancer Research UK"],"dc:creator":["Marass, Francesco"],"dc:creator.authoridentifier":["0000000289937320"],"dc:date.issued":["2017-07-22"],"dc:description.abstract":["Intratumoural heterogeneity complicates the molecular interpretation of biopsies, as multiple distinct tumour genomes are sampled and analysed at once. Ignoring the presence of these populations can lead to erroneous conclusions, and so a correct analysis must account for the clonal structure of the sample. Several methods to reconstruct tumour clonality from sequencing data have been proposed, spanning methods that either do not consider phylogenetic constraints or posit a perfect phylogeny. Models of the first type are typically latent feature models that can describe the observed data flexibly, but whose results may not be reconcilable with a phylogeny. The second type, instead, generally comprises non-parametric mixture models, with strict assumptions on the tumour’s evolutionary process. The focus of this dissertation is on the development of a phylogenetic latent feature model that can bridge the advantages of these two approaches, allowing deviations from a perfect phylogeny. The work is recounted by three statistical models of increasing complexity. First, I present a non-parametric model based on the Indian Buffet Process prior, and highlight the need for phylogenetic constraints. Second, I develop a finite, phylogenetic extension of the previous model, and show that it can outperform competing methods. Third, I generalise the phylogenetic model to arbitrary copy-number states. Markov chain Monte Carlo algorithms are presented to perform inference. The models are tested on datasets that include synthetic data, controlled biological data, and clinical data. In particular, the copy-number generalisation is applied to longitudinal circulating tumour DNA samples. Liquid biopsies that leverage circulating tumour DNA require sensitive techniques in order to detect mutations at low allele fractions. One method that allows sensitive mutation calling is the amplicon sequencing strategy TAm-Seq. I present bioinformatic tools to improve both the development of TAm-Seq amplicon panels and the analysis of its sequencing data. Finally, an enhancement of this method is presented and shown to detect mutations de novo and in a multiplexed manner at allele fractions less than 0.1%."],"dc:format.checksum.md5":["87eda9de84448d1f82354d60eee3eb5f","c4e826f60430ccf66e2129c9c012e52f"],"dc:identifier.doi":["10.17863/CAM.13715"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/dedc20c6-cc5e-466d-a7bd-61fa962c7749/download"],"dc:language":["en"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/267784"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/17fae8f7-130d-44f5-8cd2-0e7f1bd33ac2/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["Cancer genomics","Tumour heterogeneity","Computational biology","Statistical modelling","Latent feature models","Deconvolution","Amplicon sequencing"],"dc:title":["Latent feature models and non-invasive clonal reconstruction"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:03Z"}