{"id":{"repo_id":"ohiolink","oai_identifier":"oai:etd.ohiolink.edu:case1347053645"},"canonical_url":"https://search.dev.ndltd.org/etd/ohiolink/oai:etd.ohiolink.edu:case1347053645","repository":{"repo_id":"ohiolink","name":"OhioLINK","base_url":"https://etd.ohiolink.edu/acprod/odb_etd/ws/oai/oai"},"display":{"title":"Variant Detection Using Next Generation Sequencing Data","abstract":"Genetic variants, including single nucleotide polymorphisms (SNPs) and genomic structural variations (SVs), not only contribute to human diversity, but also are important to human health because some of them are proved to trigger diseases such as cancer, obesity and diabetes. With recent development of cost effective next generation sequencing (NGS) technologies, it has become possible to identify novel variants with high resolutions and to identify some copy neutral variants such as inversions, which cannot be detected using microarray based technologies. However, enormous amounts of raw sequence data generated from NGS technologies pose great challenges for data analysis. Efficient computational algorithms and tools to analyze these data are in great need. In this dissertation, we propose effective computational methods to identify genomic structural variations (deletions and inversions) and to infer accurate genotypes from multiple genotype and SNP calling algorithms using NGS. For structure variant detection, we propose a model based clustering approach utilizing a set of features defined for each type of SV events. Our method, termed SVMiner, not only provides a probability score for each candidate, but also predicts the heterozygosity of genomic deletions. Extensive experiments on genome-wide deep sequencing data have demonstrated that SVMiner is robust against the variability of a single cluster feature, and it performs well when classifying validated SV events with accentuated features. To improve SNP calling results, we propose a Na¿¿ve Bayes based approach, which combines prior information of known polymorphic sites and population specific allele frequencies with SNP calling results from multiple programs to obtain more accurate SNP genotypes. Results show that our approach has higher genotype calling accuracy than individual algorithms.","abstract_html":"Genetic variants, including single nucleotide polymorphisms (SNPs) and genomic structural variations (SVs), not only contribute to human diversity, but also are important to human health because some of them are proved to trigger diseases such as cancer, obesity and diabetes. With recent development of cost effective next generation sequencing (NGS) technologies, it has become possible to identify novel variants with high resolutions and to identify some copy neutral variants such as inversions, which cannot be detected using microarray based technologies. However, enormous amounts of raw sequence data generated from NGS technologies pose great challenges for data analysis. Efficient computational algorithms and tools to analyze these data are in great need. In this dissertation, we propose effective computational methods to identify genomic structural variations (deletions and inversions) and to infer accurate genotypes from multiple genotype and SNP calling algorithms using NGS. For structure variant detection, we propose a model based clustering approach utilizing a set of features defined for each type of SV events. Our method, termed SVMiner, not only provides a probability score for each candidate, but also predicts the heterozygosity of genomic deletions. Extensive experiments on genome-wide deep sequencing data have demonstrated that SVMiner is robust against the variability of a single cluster feature, and it performs well when classifying validated SV events with accentuated features. To improve SNP calling results, we propose a Na¿¿ve Bayes based approach, which combines prior information of known polymorphic sites and population specific allele frequencies with SNP calling results from multiple programs to obtain more accurate SNP genotypes. Results show that our approach has higher genotype calling accuracy than individual algorithms.","abstract_has_math":false,"creators":["Pyon, Yoon Soo"],"institution":"Case Western Reserve University School of Graduate Studies","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"EECS - Computer and Information Sciences","degree_department":null,"school":null,"contributors":["Li, Jing"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-03-08","date_published":"2013-03-08","updated_at":"2026-07-24T03:35:52Z","subjects":["Bioinformatics","Computer Science","Genetics","structural variation","SNP","next generation sequencing","naive Bayes classifier"],"languages":["English"],"rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://rave.ohiolink.edu/etdc/view?acc_num=case1347053645","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Li, Jing"]},{"key":"dc:creator","label":"Author","values":["Pyon, Yoon Soo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-03-08"]},{"key":"dc:publisher","label":"Institution","values":["Case Western Reserve University School of Graduate Studies / OhioLINK"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["EECS - Computer and Information Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Case Western Reserve University School of Graduate Studies"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics","Computer Science","Genetics","structural variation","SNP","next generation sequencing","naive Bayes classifier"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://rave.ohiolink.edu/etdc/view?acc_num=case1347053645"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Genetic variants, including single nucleotide polymorphisms (SNPs) and genomic structural variations (SVs), not only contribute to human diversity, but also are important to human health because some of them are proved to trigger diseases such as cancer, obesity and diabetes. With recent development of cost effective next generation sequencing (NGS) technologies, it has become possible to identify novel variants with high resolutions and to identify some copy neutral variants such as inversions, which cannot be detected using microarray based technologies. However, enormous amounts of raw sequence data generated from NGS technologies pose great challenges for data analysis. Efficient computational algorithms and tools to analyze these data are in great need. In this dissertation, we propose effective computational methods to identify genomic structural variations (deletions and inversions) and to infer accurate genotypes from multiple genotype and SNP calling algorithms using NGS. For structure variant detection, we propose a model based clustering approach utilizing a set of features defined for each type of SV events. Our method, termed SVMiner, not only provides a probability score for each candidate, but also predicts the heterozygosity of genomic deletions. Extensive experiments on genome-wide deep sequencing data have demonstrated that SVMiner is robust against the variability of a single cluster feature, and it performs well when classifying validated SV events with accentuated features. To improve SNP calling results, we propose a Na¿¿ve Bayes based approach, which combines prior information of known polymorphic sites and population specific allele frequencies with SNP calling results from multiple programs to obtain more accurate SNP genotypes. Results show that our approach has higher genotype calling accuracy than individual algorithms."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf","p.93","2.03 MB"]},{"key":"dc:title","label":"Title","values":["Variant Detection Using Next Generation Sequencing Data"]}]}],"canonical_facts":{"dc:contributor":["Li, Jing"],"dc:creator":["Pyon, Yoon Soo"],"dc:date":["2013-03-08"],"dc:description":["Genetic variants, including single nucleotide polymorphisms (SNPs) and genomic structural variations (SVs), not only contribute to human diversity, but also are important to human health because some of them are proved to trigger diseases such as cancer, obesity and diabetes. With recent development of cost effective next generation sequencing (NGS) technologies, it has become possible to identify novel variants with high resolutions and to identify some copy neutral variants such as inversions, which cannot be detected using microarray based technologies. However, enormous amounts of raw sequence data generated from NGS technologies pose great challenges for data analysis. Efficient computational algorithms and tools to analyze these data are in great need. In this dissertation, we propose effective computational methods to identify genomic structural variations (deletions and inversions) and to infer accurate genotypes from multiple genotype and SNP calling algorithms using NGS. For structure variant detection, we propose a model based clustering approach utilizing a set of features defined for each type of SV events. Our method, termed SVMiner, not only provides a probability score for each candidate, but also predicts the heterozygosity of genomic deletions. Extensive experiments on genome-wide deep sequencing data have demonstrated that SVMiner is robust against the variability of a single cluster feature, and it performs well when classifying validated SV events with accentuated features. To improve SNP calling results, we propose a Na¿¿ve Bayes based approach, which combines prior information of known polymorphic sites and population specific allele frequencies with SNP calling results from multiple programs to obtain more accurate SNP genotypes. Results show that our approach has higher genotype calling accuracy than individual algorithms."],"dc:format":["application/pdf","p.93","2.03 MB"],"dc:identifier":["http://rave.ohiolink.edu/etdc/view?acc_num=case1347053645"],"dc:language":["English"],"dc:publisher":["Case Western Reserve University School of Graduate Studies / OhioLINK"],"dc:rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"dc:subject":["Bioinformatics","Computer Science","Genetics","structural variation","SNP","next generation sequencing","naive Bayes classifier"],"dc:title":["Variant Detection Using Next Generation Sequencing Data"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["EECS - Computer and Information Sciences"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Case Western Reserve University School of Graduate Studies"]},"updated_at":"2026-07-24T03:35:52Z"}