{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/382098"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/382098","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Computational methods for the detection of somatic structural variants in cancer genomes using long-read sequencing","abstract":"Accurate detection of somatic structural variants (SVs) is critical for informing the diagnosis and treatment of human cancers. In this thesis, I present SAVANA, a computational method for the analysis of somatic SVs using long-read whole genome sequencing data from tumours and matched normal samples. SAVANA employs machine learning to distinguish true somatic SVs from germline events and noise. Additionally, I establish best practices for benchmarking SV detection using simulated and sequencing replicates to demonstrate SAVANA’s superior sensitivity, specificity, and speed compared to existing methods. I show that SAVANA performs robustly across a variety of clonality levels, genomic regions, SV types, and sizes. Using Illumina and Oxford Nanopore whole-genome sequencing data from 99 tumours and matched normal samples of patients, I show that SVs reported by SAVANA are highly concordant with those detected using short-read sequencing, including in regions of complex structural variation. I also highlight the enhanced ability of long-reads to identify SVs in repetitive regions where short-reads are unable to map with high confidence. In summary, this thesis introduces SAVANA as a novel computational method to identify somatic SVs in long-reads, establishes a robust framework for benchmarking SV detection, and demonstrates the method’s high consistency and enhanced sensitivity compared to short-reads across a large patient cohort.","abstract_html":"Accurate detection of somatic structural variants (SVs) is critical for informing the diagnosis and treatment of human cancers. In this thesis, I present SAVANA, a computational method for the analysis of somatic SVs using long-read whole genome sequencing data from tumours and matched normal samples. SAVANA employs machine learning to distinguish true somatic SVs from germline events and noise. Additionally, I establish best practices for benchmarking SV detection using simulated and sequencing replicates to demonstrate SAVANA’s superior sensitivity, specificity, and speed compared to existing methods. I show that SAVANA performs robustly across a variety of clonality levels, genomic regions, SV types, and sizes. Using Illumina and Oxford Nanopore whole-genome sequencing data from 99 tumours and matched normal samples of patients, I show that SVs reported by SAVANA are highly concordant with those detected using short-read sequencing, including in regions of complex structural variation. I also highlight the enhanced ability of long-reads to identify SVs in repetitive regions where short-reads are unable to map with high confidence. In summary, this thesis introduces SAVANA as a novel computational method to identify somatic SVs in long-reads, establishes a robust framework for benchmarking SV detection, and demonstrates the method’s high consistency and enhanced sensitivity compared to short-reads across a large patient cohort.","abstract_has_math":false,"creators":["Elrick, Hillary"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Cortes Ciriano, Isidro"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-11-04","date_published":"2024-11-04","updated_at":"2026-07-22T22:24:04Z","subjects":["bioinformatics","cancer genomics","long-read whole genome sequencing","structural variants"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d323ee04-834f-446e-ad73-7ddada068fc2/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.117082","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Cortes Ciriano, Isidro"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["EMBL-EBI"]},{"key":"dc:creator","label":"Author","values":["Elrick, Hillary"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-11-04"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/382098"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["bioinformatics","cancer genomics","long-read whole genome sequencing","structural variants"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d323ee04-834f-446e-ad73-7ddada068fc2/download","http://purl.org/NET/rdflicense/allrightsreserved"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2026-04-01"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["embargo"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.117082"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/ad5c3fb0-7fd5-448c-8976-c7ddb25c9d51/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Accurate detection of somatic structural variants (SVs) is critical for informing the diagnosis and treatment of human cancers. In this thesis, I present SAVANA, a computational method for the analysis of somatic SVs using long-read whole genome sequencing data from tumours and matched normal samples. SAVANA employs machine learning to distinguish true somatic SVs from germline events and noise. Additionally, I establish best practices for benchmarking SV detection using simulated and sequencing replicates to demonstrate SAVANA’s superior sensitivity, specificity, and speed compared to existing methods. I show that SAVANA performs robustly across a variety of clonality levels, genomic regions, SV types, and sizes. Using Illumina and Oxford Nanopore whole-genome sequencing data from 99 tumours and matched normal samples of patients, I show that SVs reported by SAVANA are highly concordant with those detected using short-read sequencing, including in regions of complex structural variation. I also highlight the enhanced ability of long-reads to identify SVs in repetitive regions where short-reads are unable to map with high confidence. In summary, this thesis introduces SAVANA as a novel computational method to identify somatic SVs in long-reads, establishes a robust framework for benchmarking SV detection, and demonstrates the method’s high consistency and enhanced sensitivity compared to short-reads across a large patient cohort."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["5f64417b757699d7e5259fae68649a96","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Computational methods for the detection of somatic structural variants in cancer genomes using long-read sequencing"]}]}],"canonical_facts":{"dc:contributor.advisor":["Cortes Ciriano, Isidro"],"dc:contributor.sponsor":["EMBL-EBI"],"dc:creator":["Elrick, Hillary"],"dc:date.issued":["2024-11-04"],"dc:description.abstract":["Accurate detection of somatic structural variants (SVs) is critical for informing the diagnosis and treatment of human cancers. In this thesis, I present SAVANA, a computational method for the analysis of somatic SVs using long-read whole genome sequencing data from tumours and matched normal samples. SAVANA employs machine learning to distinguish true somatic SVs from germline events and noise. Additionally, I establish best practices for benchmarking SV detection using simulated and sequencing replicates to demonstrate SAVANA’s superior sensitivity, specificity, and speed compared to existing methods. I show that SAVANA performs robustly across a variety of clonality levels, genomic regions, SV types, and sizes. Using Illumina and Oxford Nanopore whole-genome sequencing data from 99 tumours and matched normal samples of patients, I show that SVs reported by SAVANA are highly concordant with those detected using short-read sequencing, including in regions of complex structural variation. I also highlight the enhanced ability of long-reads to identify SVs in repetitive regions where short-reads are unable to map with high confidence. In summary, this thesis introduces SAVANA as a novel computational method to identify somatic SVs in long-reads, establishes a robust framework for benchmarking SV detection, and demonstrates the method’s high consistency and enhanced sensitivity compared to short-reads across a large patient cohort."],"dc:format.checksum.md5":["5f64417b757699d7e5259fae68649a96","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.117082"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/ad5c3fb0-7fd5-448c-8976-c7ddb25c9d51/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/382098"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d323ee04-834f-446e-ad73-7ddada068fc2/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:rights.embargodate":["2026-04-01"],"dc:rights.embargotype":["embargo"],"dc:subject":["bioinformatics","cancer genomics","long-read whole genome sequencing","structural variants"],"dc:title":["Computational methods for the detection of somatic structural variants in cancer genomes using long-read sequencing"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:04Z"}