{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/78801"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/78801","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Empirical accuracy bounds for next-generation sequencing variant calling workflows","abstract":"\"This thesis investigates the accuracy bounds imposed on alignment-based variant calling workflows due to inherent uncertainties introduced by sequencing platforms. In this work we will use simulated data to empirically quantify the maximum performance that can be expected for alignment and variant detection accuracy in a workflow. Short read sequencers are inherently incapable of producing reads that can be uniquely mapped to every position of the human reference genome, so errors are inevitable. We will analyze the repetitive content of several organisms, and estimate the maximum attainable alignment accuracy as a function of read length. Additionally, we will show that paired-end sequencing with large insert sizes (also referred to as \"\"mate-pair\"\" sequencing) is capable of mapping >99% of the human genome. We have developed a set of tools, NEAT (Next-generation Error Analysis Toolkit), which we use to create fault-injected genomic datasets. Our experiments utilize these datasets to showcase how the behavior of BWA and GATK workflows changes as a function of read lengths, error rates, quality scores, error types, and mutation types. We utilize these results to quantify the performance gains that can be expected by altering these properties of an NGS dataset. Our results highlight the sensitivity of alignment software to read lengths and error rates, and the sensitivity of variant callers to quality scores and structural variation.\"","abstract_html":"&quot;This thesis investigates the accuracy bounds imposed on alignment-based variant calling workflows due to inherent uncertainties introduced by sequencing platforms. In this work we will use simulated data to empirically quantify the maximum performance that can be expected for alignment and variant detection accuracy in a workflow. Short read sequencers are inherently incapable of producing reads that can be uniquely mapped to every position of the human reference genome, so errors are inevitable. We will analyze the repetitive content of several organisms, and estimate the maximum attainable alignment accuracy as a function of read length. Additionally, we will show that paired-end sequencing with large insert sizes (also referred to as &quot;&quot;mate-pair&quot;&quot; sequencing) is capable of mapping &gt;99% of the human genome. We have developed a set of tools, NEAT (Next-generation Error Analysis Toolkit), which we use to create fault-injected genomic datasets. Our experiments utilize these datasets to showcase how the behavior of BWA and GATK workflows changes as a function of read lengths, error rates, quality scores, error types, and mutation types. We utilize these results to quantify the performance gains that can be expected by altering these properties of an NGS dataset. Our results highlight the sensitivity of alignment software to read lengths and error rates, and the sensitivity of variant callers to quality scores and structural variation.&quot;","abstract_has_math":false,"creators":["Stephens, Zachary Daniel"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-07-22T22:46:09Z","date_published":"2015-07-22T22:46:09Z","updated_at":"2026-07-22T22:26:12Z","subjects":["Next-Generation Sequencing (NGS) Accuracy Benchmarking","Next-Generation Error Analysis Toolkit (NEAT)","Next-Generation Sequencing (NGS) Accuracy Bounds"],"languages":["en"],"rights":["Copyright 2015 Zachary Stephens"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/78801","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Stephens, Zachary Daniel"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-07-22T22:46:09Z","2017-07-23T09:15:38Z","2015-05","2015-05-01","2015-5"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Next-Generation Sequencing (NGS) Accuracy Benchmarking","Next-Generation Error Analysis Toolkit (NEAT)","Next-Generation Sequencing (NGS) Accuracy Bounds"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Zachary Stephens"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/78801"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"This thesis investigates the accuracy bounds imposed on alignment-based variant calling workflows due to inherent uncertainties introduced by sequencing platforms. In this work we will use simulated data to empirically quantify the maximum performance that can be expected for alignment and variant detection accuracy in a workflow. Short read sequencers are inherently incapable of producing reads that can be uniquely mapped to every position of the human reference genome, so errors are inevitable. We will analyze the repetitive content of several organisms, and estimate the maximum attainable alignment accuracy as a function of read length. Additionally, we will show that paired-end sequencing with large insert sizes (also referred to as \"\"mate-pair\"\" sequencing) is capable of mapping >99% of the human genome. We have developed a set of tools, NEAT (Next-generation Error Analysis Toolkit), which we use to create fault-injected genomic datasets. Our experiments utilize these datasets to showcase how the behavior of BWA and GATK workflows changes as a function of read lengths, error rates, quality scores, error types, and mutation types. We utilize these results to quantify the performance gains that can be expected by altering these properties of an NGS dataset. Our results highlight the sensitivity of alignment software to read lengths and error rates, and the sensitivity of variant callers to quality scores and structural variation.\"","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2017-05-01","The student, Zachary Stephens, accepted the attached license on 2015-04-30 at 11:28.","The student, Zachary Stephens, submitted this Thesis for approval on 2015-04-30 at 11:34.","This Thesis was approved for publication on 2015-05-01 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8210 on 2015-07-22 at 14:26:38","Made available in DSpace on 2015-07-22T22:46:09Z (GMT). No. of bitstreams: 2 STEPHENS-THESIS-2015.pdf: 2234460 bytes, checksum: b3c4cfd8c29ce0b50eee2cd1ecae3a59 (MD5) LICENSE.txt: 4213 bytes, checksum: 36132b7634d486c1ac83dadc5a9e9900 (MD5) Previous issue date: 2015-05-01","Embargo set by: Seth Robbins for item 80042 Lift date: 2017-07-22T22:46:21Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 80042 on 2017-07-23T09:15:38Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Empirical accuracy bounds for next-generation sequencing variant calling workflows"]}]}],"canonical_facts":{"dc:creator":["Stephens, Zachary Daniel"],"dc:date":["2015-07-22T22:46:09Z","2017-07-23T09:15:38Z","2015-05","2015-05-01","2015-5"],"dc:description":["\"This thesis investigates the accuracy bounds imposed on alignment-based variant calling workflows due to inherent uncertainties introduced by sequencing platforms. In this work we will use simulated data to empirically quantify the maximum performance that can be expected for alignment and variant detection accuracy in a workflow. Short read sequencers are inherently incapable of producing reads that can be uniquely mapped to every position of the human reference genome, so errors are inevitable. We will analyze the repetitive content of several organisms, and estimate the maximum attainable alignment accuracy as a function of read length. Additionally, we will show that paired-end sequencing with large insert sizes (also referred to as \"\"mate-pair\"\" sequencing) is capable of mapping >99% of the human genome. We have developed a set of tools, NEAT (Next-generation Error Analysis Toolkit), which we use to create fault-injected genomic datasets. Our experiments utilize these datasets to showcase how the behavior of BWA and GATK workflows changes as a function of read lengths, error rates, quality scores, error types, and mutation types. We utilize these results to quantify the performance gains that can be expected by altering these properties of an NGS dataset. Our results highlight the sensitivity of alignment software to read lengths and error rates, and the sensitivity of variant callers to quality scores and structural variation.\"","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2017-05-01","The student, Zachary Stephens, accepted the attached license on 2015-04-30 at 11:28.","The student, Zachary Stephens, submitted this Thesis for approval on 2015-04-30 at 11:34.","This Thesis was approved for publication on 2015-05-01 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8210 on 2015-07-22 at 14:26:38","Made available in DSpace on 2015-07-22T22:46:09Z (GMT). No. of bitstreams: 2 STEPHENS-THESIS-2015.pdf: 2234460 bytes, checksum: b3c4cfd8c29ce0b50eee2cd1ecae3a59 (MD5) LICENSE.txt: 4213 bytes, checksum: 36132b7634d486c1ac83dadc5a9e9900 (MD5) Previous issue date: 2015-05-01","Embargo set by: Seth Robbins for item 80042 Lift date: 2017-07-22T22:46:21Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 80042 on 2017-07-23T09:15:38Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/78801"],"dc:language":["en"],"dc:rights":["Copyright 2015 Zachary Stephens"],"dc:subject":["Next-Generation Sequencing (NGS) Accuracy Benchmarking","Next-Generation Error Analysis Toolkit (NEAT)","Next-Generation Sequencing (NGS) Accuracy Bounds"],"dc:title":["Empirical accuracy bounds for next-generation sequencing variant calling workflows"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:12Z"}