{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106361"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106361","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A new filtering method for improving the quality of variant discovery","abstract":"Variants identified by current genomic analysis pipelines contain many incorrectly called variants. These can be potentially eliminated by applying state-of-the-art filtering tools, such as Variant Quality Score Recalibration (VQSR) or Hard Filtering (HF). However, these methods are very user-dependent and fail to run in some cases. We propose Variant Ensemble Filter (VEF), a variant filtering tool based on decision tree ensemble methods that overcomes the main drawbacks of VQSR and HF. Contrary to these methods, we treat filtering as a supervised learning problem, using variant call data with known “true” variants, i.e., gold standard, for training. Once trained, VEF can be directly applied to filter the variants contained in a given VCF file (we consider training and testing VCF files generated with the same tools, as we assume they will share feature characteristics). For the analysis, we used Whole Genome Sequencing (WGS) human datasets for which the gold standards are available. We show on these data that the proposed filtering tool Variant Ensemble Filter (VEF) consistently outperforms VQSR and HF. In addition, we show that VEF generalizes well even when some features have missing values, when the training and testing datasets differ in coverage, and when sequencing pipelines other than GATK are used. Finally, since the training needs to be performed only once, there is a significant saving in running time when compared to VQSR (4 versus 50 minutes approximately for filtering the SNPs of a WGS Human sample).","abstract_html":"Variants identified by current genomic analysis pipelines contain many incorrectly called variants. These can be potentially eliminated by applying state-of-the-art filtering tools, such as Variant Quality Score Recalibration (VQSR) or Hard Filtering (HF). However, these methods are very user-dependent and fail to run in some cases. We propose Variant Ensemble Filter (VEF), a variant filtering tool based on decision tree ensemble methods that overcomes the main drawbacks of VQSR and HF. Contrary to these methods, we treat filtering as a supervised learning problem, using variant call data with known “true” variants, i.e., gold standard, for training. Once trained, VEF can be directly applied to filter the variants contained in a given VCF file (we consider training and testing VCF files generated with the same tools, as we assume they will share feature characteristics). For the analysis, we used Whole Genome Sequencing (WGS) human datasets for which the gold standards are available. We show on these data that the proposed filtering tool Variant Ensemble Filter (VEF) consistently outperforms VQSR and HF. In addition, we show that VEF generalizes well even when some features have missing values, when the training and testing datasets differ in coverage, and when sequencing pipelines other than GATK are used. Finally, since the training needs to be performed only once, there is a significant saving in running time when compared to VQSR (4 versus 50 minutes approximately for filtering the SNPs of a WGS Human sample).","abstract_has_math":false,"creators":["Zhang, Chuanyi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Ochoa, Idoia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T22:15:03Z","date_published":"2020-03-02T22:15:03Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Filtering","VCF file","Ensemble learning"],"languages":["en"],"rights":["Copyright 2019 Chuanyi Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106361","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ochoa, Idoia"]},{"key":"dc:creator","label":"Author","values":["Zhang, Chuanyi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T22:15:03Z","2022-03-03T10:15:22Z","2019-12-02","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Filtering","VCF file","Ensemble learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Chuanyi Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106361"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Variants identified by current genomic analysis pipelines contain many incorrectly called variants. These can be potentially eliminated by applying state-of-the-art filtering tools, such as Variant Quality Score Recalibration (VQSR) or Hard Filtering (HF). However, these methods are very user-dependent and fail to run in some cases. We propose Variant Ensemble Filter (VEF), a variant filtering tool based on decision tree ensemble methods that overcomes the main drawbacks of VQSR and HF. Contrary to these methods, we treat filtering as a supervised learning problem, using variant call data with known “true” variants, i.e., gold standard, for training. Once trained, VEF can be directly applied to filter the variants contained in a given VCF file (we consider training and testing VCF files generated with the same tools, as we assume they will share feature characteristics). For the analysis, we used Whole Genome Sequencing (WGS) human datasets for which the gold standards are available. We show on these data that the proposed filtering tool Variant Ensemble Filter (VEF) consistently outperforms VQSR and HF. In addition, we show that VEF generalizes well even when some features have missing values, when the training and testing datasets differ in coverage, and when sequencing pipelines other than GATK are used. Finally, since the training needs to be performed only once, there is a significant saving in running time when compared to VQSR (4 versus 50 minutes approximately for filtering the SNPs of a WGS Human sample).","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-12-01","The student, Chuanyi Zhang, accepted the attached license on 2019-12-02 at 09:47.","The student, Chuanyi Zhang, submitted this Thesis for approval on 2019-12-02 at 10:01.","This Thesis was approved for publication on 2019-12-02 at 10:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14645 on 2020-02-28 at 17:22:53","Made available in DSpace on 2020-03-02T22:15:03Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2019.pdf: 3572733 bytes, checksum: c4bd92ed26a1dd839a559c336d49c680 (MD5) LICENSE.txt: 4210 bytes, checksum: 00508e97af718f2a7a8d4e1b75fe7543 (MD5) Previous issue date: 2019-12-02","Embargo set by: Seth Robbins for item 113903 Lift date: 2022-03-02T22:15:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 113903 Lift date: 2022-03-02T22:18:25Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 113903 on 2022-03-03T10:15:22Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A new filtering method for improving the quality of variant discovery"]}]}],"canonical_facts":{"dc:contributor":["Ochoa, Idoia"],"dc:creator":["Zhang, Chuanyi"],"dc:date":["2020-03-02T22:15:03Z","2022-03-03T10:15:22Z","2019-12-02","2019-12"],"dc:description":["Variants identified by current genomic analysis pipelines contain many incorrectly called variants. These can be potentially eliminated by applying state-of-the-art filtering tools, such as Variant Quality Score Recalibration (VQSR) or Hard Filtering (HF). However, these methods are very user-dependent and fail to run in some cases. We propose Variant Ensemble Filter (VEF), a variant filtering tool based on decision tree ensemble methods that overcomes the main drawbacks of VQSR and HF. Contrary to these methods, we treat filtering as a supervised learning problem, using variant call data with known “true” variants, i.e., gold standard, for training. Once trained, VEF can be directly applied to filter the variants contained in a given VCF file (we consider training and testing VCF files generated with the same tools, as we assume they will share feature characteristics). For the analysis, we used Whole Genome Sequencing (WGS) human datasets for which the gold standards are available. We show on these data that the proposed filtering tool Variant Ensemble Filter (VEF) consistently outperforms VQSR and HF. In addition, we show that VEF generalizes well even when some features have missing values, when the training and testing datasets differ in coverage, and when sequencing pipelines other than GATK are used. Finally, since the training needs to be performed only once, there is a significant saving in running time when compared to VQSR (4 versus 50 minutes approximately for filtering the SNPs of a WGS Human sample).","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-12-01","The student, Chuanyi Zhang, accepted the attached license on 2019-12-02 at 09:47.","The student, Chuanyi Zhang, submitted this Thesis for approval on 2019-12-02 at 10:01.","This Thesis was approved for publication on 2019-12-02 at 10:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14645 on 2020-02-28 at 17:22:53","Made available in DSpace on 2020-03-02T22:15:03Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2019.pdf: 3572733 bytes, checksum: c4bd92ed26a1dd839a559c336d49c680 (MD5) LICENSE.txt: 4210 bytes, checksum: 00508e97af718f2a7a8d4e1b75fe7543 (MD5) Previous issue date: 2019-12-02","Embargo set by: Seth Robbins for item 113903 Lift date: 2022-03-02T22:15:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 113903 Lift date: 2022-03-02T22:18:25Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 113903 on 2022-03-03T10:15:22Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106361"],"dc:language":["en"],"dc:rights":["Copyright 2019 Chuanyi Zhang"],"dc:subject":["Filtering","VCF file","Ensemble learning"],"dc:title":["A new filtering method for improving the quality of variant discovery"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}