{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/109337"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/109337","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Navigating through the uncertainty of genotyping-by-sequencing data in polyploids","abstract":"The development of genotyping-by-sequencing (GBS) methods has facilitated genomics studies in non-model species, including polyploids. Variant and genotype calling methods have been established for autopolyploids but for a species with a complex genome, such as sugarcane, the level of uncertainty within GBS data increases making trait mapping difficult. Furthermore, variant and genotype calling methods remain a challenge for both recent and ancient allopolyploids (e.g. wheat, maize, soybean, Miscanthus), particularly where the reference genome contains highly similar paralogous sequences that do not pair at meiosis. Alignment of sequence tags to the appropriate position within highly duplicated reference genomes remains a challenge inadequately addressed by existing alignment software. Although some variant calling pipelines can discriminate a paralogous locus from a Mendelian locus, the detection of these paralogous loci is typically for the purpose of the exclusion of these loci from the downstream analysis of genomic studies. We explore the significance of eliminating paralogous loci in downstream analysis using a newly developed pipeline developed to sort sequence tags to their correct alignment locations based on the novel Hind/HE statistic. The goal of this study was to evaluate the sorting pipeline’s ability to properly align paralogous loci to the correct position with respect to the reference genome. Three studies were conducted with a population of 400 individuals simulated based upon the Triticum aestivum, the reanalysis of a previously published genome-wide study of fusarium head blight in 273 wheat breeding lines, and the reanalysis of a previously published genome-wide study of traits associated with yield in a Miscanthus diversity panel. Results from the study suggested that the filtering of sequences using the Hind/HE statistic underlying polyRAD v1.2 may lead differences in the output of sequences. Further comparison of each output suggested that the output of the novel pipeline, polyRAD, was concentrated in gene-rich regions compared to other standard variant calling pipelines. From this study, we provide recommendations for future users of the polyRAD v1.2 variant calling pipeline. Overall we recommend that polyRAD v1.2 is more useful for populations of outcrossing species.","abstract_html":"The development of genotyping-by-sequencing (GBS) methods has facilitated genomics studies in non-model species, including polyploids. Variant and genotype calling methods have been established for autopolyploids but for a species with a complex genome, such as sugarcane, the level of uncertainty within GBS data increases making trait mapping difficult. Furthermore, variant and genotype calling methods remain a challenge for both recent and ancient allopolyploids (e.g. wheat, maize, soybean, Miscanthus), particularly where the reference genome contains highly similar paralogous sequences that do not pair at meiosis. Alignment of sequence tags to the appropriate position within highly duplicated reference genomes remains a challenge inadequately addressed by existing alignment software. Although some variant calling pipelines can discriminate a paralogous locus from a Mendelian locus, the detection of these paralogous loci is typically for the purpose of the exclusion of these loci from the downstream analysis of genomic studies. We explore the significance of eliminating paralogous loci in downstream analysis using a newly developed pipeline developed to sort sequence tags to their correct alignment locations based on the novel Hind/HE statistic. The goal of this study was to evaluate the sorting pipeline’s ability to properly align paralogous loci to the correct position with respect to the reference genome. Three studies were conducted with a population of 400 individuals simulated based upon the Triticum aestivum, the reanalysis of a previously published genome-wide study of fusarium head blight in 273 wheat breeding lines, and the reanalysis of a previously published genome-wide study of traits associated with yield in a Miscanthus diversity panel. Results from the study suggested that the filtering of sequences using the Hind/HE statistic underlying polyRAD v1.2 may lead differences in the output of sequences. Further comparison of each output suggested that the output of the novel pipeline, polyRAD, was concentrated in gene-rich regions compared to other standard variant calling pipelines. From this study, we provide recommendations for future users of the polyRAD v1.2 variant calling pipeline. Overall we recommend that polyRAD v1.2 is more useful for populations of outcrossing species.","abstract_has_math":false,"creators":["Mays, Wittney Debora"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Bioinformatics","degree_department":null,"school":null,"contributors":["Sacks, Erik J","Clark, Lindsay V","Ming, Ray","Lipka, Alexander E"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-03-05T21:36:48Z","date_published":"2021-03-05T21:36:48Z","updated_at":"2026-07-22T22:24:50Z","subjects":["bioinformatics","variant calling","polyploidy","genotyping-by-sequencing","GBS"],"languages":["en"],"rights":["Copyright 2020 Wittney Mays"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/109337","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sacks, Erik J","Clark, Lindsay V","Ming, Ray","Lipka, Alexander E"]},{"key":"dc:creator","label":"Author","values":["Mays, Wittney Debora"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-03-05T21:36:48Z","2020-10-05","2020-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Bioinformatics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["bioinformatics","variant calling","polyploidy","genotyping-by-sequencing","GBS"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Wittney Mays"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/109337"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The development of genotyping-by-sequencing (GBS) methods has facilitated genomics studies in non-model species, including polyploids. Variant and genotype calling methods have been established for autopolyploids but for a species with a complex genome, such as sugarcane, the level of uncertainty within GBS data increases making trait mapping difficult. Furthermore, variant and genotype calling methods remain a challenge for both recent and ancient allopolyploids (e.g. wheat, maize, soybean, Miscanthus), particularly where the reference genome contains highly similar paralogous sequences that do not pair at meiosis. Alignment of sequence tags to the appropriate position within highly duplicated reference genomes remains a challenge inadequately addressed by existing alignment software. Although some variant calling pipelines can discriminate a paralogous locus from a Mendelian locus, the detection of these paralogous loci is typically for the purpose of the exclusion of these loci from the downstream analysis of genomic studies. We explore the significance of eliminating paralogous loci in downstream analysis using a newly developed pipeline developed to sort sequence tags to their correct alignment locations based on the novel Hind/HE statistic. The goal of this study was to evaluate the sorting pipeline’s ability to properly align paralogous loci to the correct position with respect to the reference genome. Three studies were conducted with a population of 400 individuals simulated based upon the Triticum aestivum, the reanalysis of a previously published genome-wide study of fusarium head blight in 273 wheat breeding lines, and the reanalysis of a previously published genome-wide study of traits associated with yield in a Miscanthus diversity panel. Results from the study suggested that the filtering of sequences using the Hind/HE statistic underlying polyRAD v1.2 may lead differences in the output of sequences. Further comparison of each output suggested that the output of the novel pipeline, polyRAD, was concentrated in gene-rich regions compared to other standard variant calling pipelines. From this study, we provide recommendations for future users of the polyRAD v1.2 variant calling pipeline. Overall we recommend that polyRAD v1.2 is more useful for populations of outcrossing species.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Wittney Mays, accepted the attached license on 2020-09-13 at 01:50.","The student, Wittney Mays, submitted this Thesis for approval on 2020-09-13 at 01:57.","This Thesis was approved for publication on 2020-10-05 at 09:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15809 on 2021-03-04 at 15:33:50","Made available in DSpace on 2021-03-05T21:36:48Z (GMT). No. of bitstreams: 2 MAYS-THESIS-2020.pdf: 2735820 bytes, checksum: 88277f856f5d979c1ed557a2442e99c5 (MD5) LICENSE.txt: 4209 bytes, checksum: 8590c14b394620511b5254076248c917 (MD5) Previous issue date: 2020-10-05"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Navigating through the uncertainty of genotyping-by-sequencing data in polyploids"]}]}],"canonical_facts":{"dc:contributor":["Sacks, Erik J","Clark, Lindsay V","Ming, Ray","Lipka, Alexander E"],"dc:creator":["Mays, Wittney Debora"],"dc:date":["2021-03-05T21:36:48Z","2020-10-05","2020-12"],"dc:description":["The development of genotyping-by-sequencing (GBS) methods has facilitated genomics studies in non-model species, including polyploids. Variant and genotype calling methods have been established for autopolyploids but for a species with a complex genome, such as sugarcane, the level of uncertainty within GBS data increases making trait mapping difficult. Furthermore, variant and genotype calling methods remain a challenge for both recent and ancient allopolyploids (e.g. wheat, maize, soybean, Miscanthus), particularly where the reference genome contains highly similar paralogous sequences that do not pair at meiosis. Alignment of sequence tags to the appropriate position within highly duplicated reference genomes remains a challenge inadequately addressed by existing alignment software. Although some variant calling pipelines can discriminate a paralogous locus from a Mendelian locus, the detection of these paralogous loci is typically for the purpose of the exclusion of these loci from the downstream analysis of genomic studies. We explore the significance of eliminating paralogous loci in downstream analysis using a newly developed pipeline developed to sort sequence tags to their correct alignment locations based on the novel Hind/HE statistic. The goal of this study was to evaluate the sorting pipeline’s ability to properly align paralogous loci to the correct position with respect to the reference genome. Three studies were conducted with a population of 400 individuals simulated based upon the Triticum aestivum, the reanalysis of a previously published genome-wide study of fusarium head blight in 273 wheat breeding lines, and the reanalysis of a previously published genome-wide study of traits associated with yield in a Miscanthus diversity panel. Results from the study suggested that the filtering of sequences using the Hind/HE statistic underlying polyRAD v1.2 may lead differences in the output of sequences. Further comparison of each output suggested that the output of the novel pipeline, polyRAD, was concentrated in gene-rich regions compared to other standard variant calling pipelines. From this study, we provide recommendations for future users of the polyRAD v1.2 variant calling pipeline. Overall we recommend that polyRAD v1.2 is more useful for populations of outcrossing species.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Wittney Mays, accepted the attached license on 2020-09-13 at 01:50.","The student, Wittney Mays, submitted this Thesis for approval on 2020-09-13 at 01:57.","This Thesis was approved for publication on 2020-10-05 at 09:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15809 on 2021-03-04 at 15:33:50","Made available in DSpace on 2021-03-05T21:36:48Z (GMT). No. of bitstreams: 2 MAYS-THESIS-2020.pdf: 2735820 bytes, checksum: 88277f856f5d979c1ed557a2442e99c5 (MD5) LICENSE.txt: 4209 bytes, checksum: 8590c14b394620511b5254076248c917 (MD5) Previous issue date: 2020-10-05"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/109337"],"dc:language":["en"],"dc:rights":["Copyright 2020 Wittney Mays"],"dc:subject":["bioinformatics","variant calling","polyploidy","genotyping-by-sequencing","GBS"],"dc:title":["Navigating through the uncertainty of genotyping-by-sequencing data in polyploids"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Bioinformatics"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}