{"id":{"repo_id":"dundee","oai_identifier":"oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"},"canonical_url":"https://search.dev.ndltd.org/etd/dundee/oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47","repository":{"repo_id":"dundee","name":"University of Dundee","base_url":"https://discovery.dundee.ac.uk/ws/oai"},"display":{"title":"A Quantitative Exploration of Causes of False Positive Single Nucleotide Polymorphisms in Next-Generation Sequencing Data","abstract":"Single Nucleotide Polymorphisms (SNPs) are widely used molecular markers, and their use has increased massively since the inception of Next-Generation Sequencing (NGS) technologies, which allow detection of large numbers of SNPs at low cost. However, both NGS data and their analysis are error-prone, which can lead to the generation of false positive (FP) SNPs. The traditional approach to SNP discovery is based on mapping reads to a reference sequence. Apart from sequencing errors, which vary in pattern and rate depending on the sequencing platform, the short read lengths that prevail in NGS, together with the repetitive nature of the genomes of many organisms, can lead to errors in the genome assembly and/or read mapping stages of the mapping-based approach for SNP discovery.<br/><br/>The work described here has investigated and quantified some mechanisms that cause false positive SNPs. These include reference misassembly due to the presence of paralogous sequences and read cross-mapping, along with associated factors such as quality of the reference sequence, read length, choice of mapper and variant caller, mapping stringency, and filtering of SNPs by read mapping quality and read depth. The study shows that both paralogs and the choice of tools and parameters involved in variant calling can have a dramatic effect on the number of FP SNPs produced. A brief exploration of the influence of these factors towards false negative (FN) SNPs generation is also carried out in the end of the study, paving the way to new insights. This thesis aims to provide a stepping stone towards a better understanding of the factors influencing the mapping-based SNP discovery approach.","abstract_html":"Single Nucleotide Polymorphisms (SNPs) are widely used molecular markers, and their use has increased massively since the inception of Next-Generation Sequencing (NGS) technologies, which allow detection of large numbers of SNPs at low cost. However, both NGS data and their analysis are error-prone, which can lead to the generation of false positive (FP) SNPs. The traditional approach to SNP discovery is based on mapping reads to a reference sequence. Apart from sequencing errors, which vary in pattern and rate depending on the sequencing platform, the short read lengths that prevail in NGS, together with the repetitive nature of the genomes of many organisms, can lead to errors in the genome assembly and/or read mapping stages of the mapping-based approach for SNP discovery.&lt;br/&gt;&lt;br/&gt;The work described here has investigated and quantified some mechanisms that cause false positive SNPs. These include reference misassembly due to the presence of paralogous sequences and read cross-mapping, along with associated factors such as quality of the reference sequence, read length, choice of mapper and variant caller, mapping stringency, and filtering of SNPs by read mapping quality and read depth. The study shows that both paralogs and the choice of tools and parameters involved in variant calling can have a dramatic effect on the number of FP SNPs produced. A brief exploration of the influence of these factors towards false negative (FN) SNPs generation is also carried out in the end of the study, paving the way to new insights. This thesis aims to provide a stepping stone towards a better understanding of the factors influencing the mapping-based SNP discovery approach.","abstract_has_math":false,"creators":["Bello Ribeiro, Antonio Claudio"],"institution":"University of Dundee","degree_name":"Doctor of Philosophy","degree_level":"Doctoral Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016","date_published":"2016","updated_at":"2026-07-24T02:08:19Z","subjects":["False positive SNP","NGS","Read mismapping","Misassembly","Mapping stringency","Read lengths"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"],"render_values":[{"text":"oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47","href":null,"code":true}]}]},"links":{"outbound_url":"https://discovery.dundee.ac.uk/en/studentTheses/7be82ce3-7263-4d6f-880f-71d3fceeed47","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.sponsor","label":"Sponsor","values":["The James Hutton Institute"]},{"key":"dc:creator","label":"Author","values":["Bello Ribeiro, Antonio Claudio"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016"]},{"key":"dc:date.issued","label":"Date","values":["2016"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Plant Sciences"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Dundee"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://discovery.dundee.ac.uk/en/studentTheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral Thesis"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["False positive SNP","NGS","Read mismapping","Misassembly","Mapping stringency","Read lengths"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2017-10-31"]},{"key":"dc:rights.embargoreason","label":"Dc Rights Embargoreason","values":["/dk/atira/pure/core/document/studentthesisembargoreason/commercialexploitation"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47","https://discovery.dundee.ac.uk/en/studentTheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://discovery.dundee.ac.uk/files/36512126/20161031_AntonioCBRibeiro_Thesis_UoD_JHI.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Single Nucleotide Polymorphisms (SNPs) are widely used molecular markers, and their use has increased massively since the inception of Next-Generation Sequencing (NGS) technologies, which allow detection of large numbers of SNPs at low cost. However, both NGS data and their analysis are error-prone, which can lead to the generation of false positive (FP) SNPs. The traditional approach to SNP discovery is based on mapping reads to a reference sequence. Apart from sequencing errors, which vary in pattern and rate depending on the sequencing platform, the short read lengths that prevail in NGS, together with the repetitive nature of the genomes of many organisms, can lead to errors in the genome assembly and/or read mapping stages of the mapping-based approach for SNP discovery.<br/><br/>The work described here has investigated and quantified some mechanisms that cause false positive SNPs. These include reference misassembly due to the presence of paralogous sequences and read cross-mapping, along with associated factors such as quality of the reference sequence, read length, choice of mapper and variant caller, mapping stringency, and filtering of SNPs by read mapping quality and read depth. The study shows that both paralogs and the choice of tools and parameters involved in variant calling can have a dramatic effect on the number of FP SNPs produced. A brief exploration of the influence of these factors towards false negative (FN) SNPs generation is also carried out in the end of the study, paving the way to new insights. This thesis aims to provide a stepping stone towards a better understanding of the factors influencing the mapping-based SNP discovery approach."]},{"key":"dc:title","label":"Title","values":["A Quantitative Exploration of Causes of False Positive Single Nucleotide Polymorphisms in Next-Generation Sequencing Data"]}]}],"canonical_facts":{"dc:contributor.sponsor":["The James Hutton Institute"],"dc:creator":["Bello Ribeiro, Antonio Claudio"],"dc:date":["2016"],"dc:date.issued":["2016"],"dc:description.abstract":["Single Nucleotide Polymorphisms (SNPs) are widely used molecular markers, and their use has increased massively since the inception of Next-Generation Sequencing (NGS) technologies, which allow detection of large numbers of SNPs at low cost. However, both NGS data and their analysis are error-prone, which can lead to the generation of false positive (FP) SNPs. The traditional approach to SNP discovery is based on mapping reads to a reference sequence. Apart from sequencing errors, which vary in pattern and rate depending on the sequencing platform, the short read lengths that prevail in NGS, together with the repetitive nature of the genomes of many organisms, can lead to errors in the genome assembly and/or read mapping stages of the mapping-based approach for SNP discovery.<br/><br/>The work described here has investigated and quantified some mechanisms that cause false positive SNPs. These include reference misassembly due to the presence of paralogous sequences and read cross-mapping, along with associated factors such as quality of the reference sequence, read length, choice of mapper and variant caller, mapping stringency, and filtering of SNPs by read mapping quality and read depth. The study shows that both paralogs and the choice of tools and parameters involved in variant calling can have a dramatic effect on the number of FP SNPs produced. A brief exploration of the influence of these factors towards false negative (FN) SNPs generation is also carried out in the end of the study, paving the way to new insights. This thesis aims to provide a stepping stone towards a better understanding of the factors influencing the mapping-based SNP discovery approach."],"dc:identifier":["oai:discovery.dundee.ac.uk:studenttheses/7be82ce3-7263-4d6f-880f-71d3fceeed47","https://discovery.dundee.ac.uk/en/studentTheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"],"dc:identifier.uri":["https://discovery.dundee.ac.uk/files/36512126/20161031_AntonioCBRibeiro_Thesis_UoD_JHI.pdf"],"dc:language":["eng"],"dc:publisher.department":["Plant Sciences"],"dc:publisher.institution":["University of Dundee"],"dc:relation.isreferencedby":["https://discovery.dundee.ac.uk/en/studentTheses/7be82ce3-7263-4d6f-880f-71d3fceeed47"],"dc:rights.embargodate":["2017-10-31"],"dc:rights.embargoreason":["/dk/atira/pure/core/document/studentthesisembargoreason/commercialexploitation"],"dc:subject":["False positive SNP","NGS","Read mismapping","Misassembly","Mapping stringency","Read lengths"],"dc:title":["A Quantitative Exploration of Causes of False Positive Single Nucleotide Polymorphisms in Next-Generation Sequencing Data"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral Thesis"],"dc:type.qualificationname":["Doctor of Philosophy"]},"updated_at":"2026-07-24T02:08:19Z"}