{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/304029"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/304029","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Variation-aware algorithms for cancer genome analysis","abstract":"In this thesis, I explore variation-aware algorithms for analyzing cancer genomes. The scientific community has extensively catalogued millions of mutations present in cancer cells. This information is rarely used during read alignment and variant calling because of a lack of algorithms for doing so. Rediscovering these variants wastes significant computational time and negatively impacts the sensitivity of detection, motivating the development of new solutions. The variants in malignant cells can arise from a number of genetic and environmental sources. In the years after the Chernobyl nuclear disaster, thousands of individuals in areas where radionuclides were deposited developed thyroid cancer. The excess relative risk of thyroid cancer has been estimated to be between fifteen and thirty fold higher following $^{131}I$ exposure. In Chapter 2, I analyze more than 300 thyroid cancer cases from children and young adults exposed to ionizing radiation from increased levels of $^{131}I$ originating from Chernobyl. I characterize the mutational landscape of these tumors across the variant size spectrum and compare it to sporadic thyroid cancer cases. I investigate possible signs of radiation exposure in the genome, especially large balanced structural variants and small indels. In Chapter 3, I develop methods for working with structural variants in variation graphs. Variation graphs have been shown to reduce reference bias and improve alignment to variant sites, though previous work has primarily focused on variants less than fifty base pairs in size. I describe the tradeoffs of various representations of large variants in the graph. I then develop several methods for genotyping and calling large variants in graphs. I construct graphs of both germline and somatic variation and describe how to work with these structures. I develop a method for fast structural variant genotyping using graph mappings and show that this significantly outperforms standard structural variant callers for certain types of variation. I describe how to locate structural variant mismapping signatures on variation graphs and how variation graphs can improve calling of structural variants. Lastly in Chapter 4, I demonstrate a new application of an alignment-free algorithm for genome analysis. I describe a MinHash toolkit for viral coinfection analysis and its application to Human Papillomavirus (HPV) samples. This toolkit is able to classify individual reads from multiple sequencing technologies and accurately detect clinically- relevant HPV coinfections. Finally, I dicuss how these approaches can be applied in other genomic analyses.","abstract_html":"In this thesis, I explore variation-aware algorithms for analyzing cancer genomes. The scientific community has extensively catalogued millions of mutations present in cancer cells. This information is rarely used during read alignment and variant calling because of a lack of algorithms for doing so. Rediscovering these variants wastes significant computational time and negatively impacts the sensitivity of detection, motivating the development of new solutions. The variants in malignant cells can arise from a number of genetic and environmental sources. In the years after the Chernobyl nuclear disaster, thousands of individuals in areas where radionuclides were deposited developed thyroid cancer. The excess relative risk of thyroid cancer has been estimated to be between fifteen and thirty fold higher following <span class=\"etd-inline-math\"><sup>131</sup>I</span> exposure. In Chapter 2, I analyze more than 300 thyroid cancer cases from children and young adults exposed to ionizing radiation from increased levels of <span class=\"etd-inline-math\"><sup>131</sup>I</span> originating from Chernobyl. I characterize the mutational landscape of these tumors across the variant size spectrum and compare it to sporadic thyroid cancer cases. I investigate possible signs of radiation exposure in the genome, especially large balanced structural variants and small indels. In Chapter 3, I develop methods for working with structural variants in variation graphs. Variation graphs have been shown to reduce reference bias and improve alignment to variant sites, though previous work has primarily focused on variants less than fifty base pairs in size. I describe the tradeoffs of various representations of large variants in the graph. I then develop several methods for genotyping and calling large variants in graphs. I construct graphs of both germline and somatic variation and describe how to work with these structures. I develop a method for fast structural variant genotyping using graph mappings and show that this significantly outperforms standard structural variant callers for certain types of variation. I describe how to locate structural variant mismapping signatures on variation graphs and how variation graphs can improve calling of structural variants. Lastly in Chapter 4, I demonstrate a new application of an alignment-free algorithm for genome analysis. I describe a MinHash toolkit for viral coinfection analysis and its application to Human Papillomavirus (HPV) samples. This toolkit is able to classify individual reads from multiple sequencing technologies and accurately detect clinically- relevant HPV coinfections. Finally, I dicuss how these approaches can be applied in other genomic analyses.","abstract_has_math":true,"creators":["Dawson, Eric Thomas"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Durbin, Richard","Chanock, Stephen"],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-05-16","date_published":"2020-05-16","updated_at":"2026-07-22T22:23:56Z","subjects":["Cancer","Genomics","Bioinformatics","Variation Graphs","Structural Variation","Ionizing Radiation","Chernobyl","Algorithms","Kmers","Human Papillomavirus","HPV","MinHash","Thyroid Cancer","Papillary Thyroid Carcinoma","Mutational Signatures","Indels","Genetics","Genetic Variation","Variant Calling","Genome Assembly","Somatic Variation"],"languages":["en"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/6005c3b2-1b6c-47b4-b1df-3b5535887667/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000154481653","0000000291301006","0000000223243393"],"render_values":[{"text":"0000-0001-5448-1653","href":"https://orcid.org/0000-0001-5448-1653","code":true},{"text":"0000-0002-9130-1006","href":"https://orcid.org/0000-0002-9130-1006","code":true},{"text":"0000-0002-2324-3393","href":"https://orcid.org/0000-0002-2324-3393","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.51110","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Durbin, Richard","Chanock, Stephen"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["The majority of this thesis was supported by funds from the US National Institutes of Health Intramural Research Program as well as the Wellcome Sanger Institute. ETD was also supported by an NIH - Cambridge Trust Fellowship."]},{"key":"dc:creator","label":"Author","values":["Dawson, Eric Thomas"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000154481653","0000000291301006","0000000223243393"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2020-05-16"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/304029"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cancer","Genomics","Bioinformatics","Variation Graphs","Structural Variation","Ionizing Radiation","Chernobyl","Algorithms","Kmers","Human Papillomavirus","HPV","MinHash","Thyroid Cancer","Papillary Thyroid Carcinoma","Mutational Signatures","Indels","Genetics","Genetic Variation","Variant Calling","Genome Assembly","Somatic Variation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/6005c3b2-1b6c-47b4-b1df-3b5535887667/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.17863/CAM.51110"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/b590e764-05e8-4e2b-a652-6e99c113e7eb/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, I explore variation-aware algorithms for analyzing cancer genomes. The scientific community has extensively catalogued millions of mutations present in cancer cells. This information is rarely used during read alignment and variant calling because of a lack of algorithms for doing so. Rediscovering these variants wastes significant computational time and negatively impacts the sensitivity of detection, motivating the development of new solutions. The variants in malignant cells can arise from a number of genetic and environmental sources. In the years after the Chernobyl nuclear disaster, thousands of individuals in areas where radionuclides were deposited developed thyroid cancer. The excess relative risk of thyroid cancer has been estimated to be between fifteen and thirty fold higher following $^{131}I$ exposure. In Chapter 2, I analyze more than 300 thyroid cancer cases from children and young adults exposed to ionizing radiation from increased levels of $^{131}I$ originating from Chernobyl. I characterize the mutational landscape of these tumors across the variant size spectrum and compare it to sporadic thyroid cancer cases. I investigate possible signs of radiation exposure in the genome, especially large balanced structural variants and small indels. In Chapter 3, I develop methods for working with structural variants in variation graphs. Variation graphs have been shown to reduce reference bias and improve alignment to variant sites, though previous work has primarily focused on variants less than fifty base pairs in size. I describe the tradeoffs of various representations of large variants in the graph. I then develop several methods for genotyping and calling large variants in graphs. I construct graphs of both germline and somatic variation and describe how to work with these structures. I develop a method for fast structural variant genotyping using graph mappings and show that this significantly outperforms standard structural variant callers for certain types of variation. I describe how to locate structural variant mismapping signatures on variation graphs and how variation graphs can improve calling of structural variants. Lastly in Chapter 4, I demonstrate a new application of an alignment-free algorithm for genome analysis. I describe a MinHash toolkit for viral coinfection analysis and its application to Human Papillomavirus (HPV) samples. This toolkit is able to classify individual reads from multiple sequencing technologies and accurately detect clinically- relevant HPV coinfections. Finally, I dicuss how these approaches can be applied in other genomic analyses."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["a859d91548a767475e3815c98f382edc","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Variation-aware algorithms for cancer genome analysis"]}]}],"canonical_facts":{"dc:contributor.advisor":["Durbin, Richard","Chanock, Stephen"],"dc:contributor.sponsor":["The majority of this thesis was supported by funds from the US National Institutes of Health Intramural Research Program as well as the Wellcome Sanger Institute. ETD was also supported by an NIH - Cambridge Trust Fellowship."],"dc:creator":["Dawson, Eric Thomas"],"dc:creator.authoridentifier":["0000000154481653","0000000291301006","0000000223243393"],"dc:date.issued":["2020-05-16"],"dc:description.abstract":["In this thesis, I explore variation-aware algorithms for analyzing cancer genomes. The scientific community has extensively catalogued millions of mutations present in cancer cells. This information is rarely used during read alignment and variant calling because of a lack of algorithms for doing so. Rediscovering these variants wastes significant computational time and negatively impacts the sensitivity of detection, motivating the development of new solutions. The variants in malignant cells can arise from a number of genetic and environmental sources. In the years after the Chernobyl nuclear disaster, thousands of individuals in areas where radionuclides were deposited developed thyroid cancer. The excess relative risk of thyroid cancer has been estimated to be between fifteen and thirty fold higher following $^{131}I$ exposure. In Chapter 2, I analyze more than 300 thyroid cancer cases from children and young adults exposed to ionizing radiation from increased levels of $^{131}I$ originating from Chernobyl. I characterize the mutational landscape of these tumors across the variant size spectrum and compare it to sporadic thyroid cancer cases. I investigate possible signs of radiation exposure in the genome, especially large balanced structural variants and small indels. In Chapter 3, I develop methods for working with structural variants in variation graphs. Variation graphs have been shown to reduce reference bias and improve alignment to variant sites, though previous work has primarily focused on variants less than fifty base pairs in size. I describe the tradeoffs of various representations of large variants in the graph. I then develop several methods for genotyping and calling large variants in graphs. I construct graphs of both germline and somatic variation and describe how to work with these structures. I develop a method for fast structural variant genotyping using graph mappings and show that this significantly outperforms standard structural variant callers for certain types of variation. I describe how to locate structural variant mismapping signatures on variation graphs and how variation graphs can improve calling of structural variants. Lastly in Chapter 4, I demonstrate a new application of an alignment-free algorithm for genome analysis. I describe a MinHash toolkit for viral coinfection analysis and its application to Human Papillomavirus (HPV) samples. This toolkit is able to classify individual reads from multiple sequencing technologies and accurately detect clinically- relevant HPV coinfections. Finally, I dicuss how these approaches can be applied in other genomic analyses."],"dc:format.checksum.md5":["a859d91548a767475e3815c98f382edc","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["10.17863/CAM.51110"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/b590e764-05e8-4e2b-a652-6e99c113e7eb/download"],"dc:language":["en"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/304029"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/6005c3b2-1b6c-47b4-b1df-3b5535887667/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Cancer","Genomics","Bioinformatics","Variation Graphs","Structural Variation","Ionizing Radiation","Chernobyl","Algorithms","Kmers","Human Papillomavirus","HPV","MinHash","Thyroid Cancer","Papillary Thyroid Carcinoma","Mutational Signatures","Indels","Genetics","Genetic Variation","Variant Calling","Genome Assembly","Somatic Variation"],"dc:title":["Variation-aware algorithms for cancer genome analysis"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:56Z"}