{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106182"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106182","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large scale phylogenomic estimation","abstract":"Phylogenomic estimation - the science of calculating evolutionary trees from genomic data - is an important biological problem. As the amount of genomic data in biological datasets increases, new methods are needed to analyze this data. Cutting edge analyses may utilize genomes from tens of thousands of species. I present several methods for supertree and species tree estimation: ASTRID, FastRFS, SVDquest, and SIESTA. ASTRID can be used for both species tree and supertree estimation, and is designed to scale to very large datasets while maintaining a high level of accuracy. FastRFS is a supertree method that uses an exact constrained optimization algorithm to find accurate supertrees. SVDquest is a coalescent-aware species tree estimation method that estimates trees directly from sequences without using gene trees. Finally, SIESTA is a modification to the algorithms used by FastRFS, SVDquest, and other methods including ASTRAL that allows for the output and analysis of multiple optimal solutions, if they exist. For all these methods, I describe the algorithms used, along with a theoretical analysis of their running time and their statistical consistency. I also show results on biological and simulated data that demonstrate these methods’ effectiveness over a wide range of model conditions. In addition, I present the results of an experiment that compares various methods on trees simulated under both incomplete lineage sorting (ILS) as well as horizontal gene transfer (HGT).","abstract_html":"Phylogenomic estimation - the science of calculating evolutionary trees from genomic data - is an important biological problem. As the amount of genomic data in biological datasets increases, new methods are needed to analyze this data. Cutting edge analyses may utilize genomes from tens of thousands of species. I present several methods for supertree and species tree estimation: ASTRID, FastRFS, SVDquest, and SIESTA. ASTRID can be used for both species tree and supertree estimation, and is designed to scale to very large datasets while maintaining a high level of accuracy. FastRFS is a supertree method that uses an exact constrained optimization algorithm to find accurate supertrees. SVDquest is a coalescent-aware species tree estimation method that estimates trees directly from sequences without using gene trees. Finally, SIESTA is a modification to the algorithms used by FastRFS, SVDquest, and other methods including ASTRAL that allows for the output and analysis of multiple optimal solutions, if they exist. For all these methods, I describe the algorithms used, along with a theoretical analysis of their running time and their statistical consistency. I also show results on biological and simulated data that demonstrate these methods’ effectiveness over a wide range of model conditions. In addition, I present the results of an experiment that compares various methods on trees simulated under both incomplete lineage sorting (ILS) as well as horizontal gene transfer (HGT).","abstract_has_math":false,"creators":["Vachaspati, Pranjal"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Warnow, Tandy","Amato, Nancy","Chekuri, Chandra","Leebens-Mack, James"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T21:58:07Z","date_published":"2020-03-02T21:58:07Z","updated_at":"2026-07-22T22:24:45Z","subjects":["phylogenomics","species trees","computational biology","phylogenetics","evolution","incomplete lineage sorting","multispecies coalescent"],"languages":["en"],"rights":["Copyright 2019 Pranjal Vachaspati"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106182","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Warnow, Tandy","Amato, Nancy","Chekuri, Chandra","Leebens-Mack, James"]},{"key":"dc:creator","label":"Author","values":["Vachaspati, Pranjal"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T21:58:07Z","2019-12-01","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["phylogenomics","species trees","computational biology","phylogenetics","evolution","incomplete lineage sorting","multispecies coalescent"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Pranjal Vachaspati"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106182"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Phylogenomic estimation - the science of calculating evolutionary trees from genomic data - is an important biological problem. As the amount of genomic data in biological datasets increases, new methods are needed to analyze this data. Cutting edge analyses may utilize genomes from tens of thousands of species. I present several methods for supertree and species tree estimation: ASTRID, FastRFS, SVDquest, and SIESTA. ASTRID can be used for both species tree and supertree estimation, and is designed to scale to very large datasets while maintaining a high level of accuracy. FastRFS is a supertree method that uses an exact constrained optimization algorithm to find accurate supertrees. SVDquest is a coalescent-aware species tree estimation method that estimates trees directly from sequences without using gene trees. Finally, SIESTA is a modification to the algorithms used by FastRFS, SVDquest, and other methods including ASTRAL that allows for the output and analysis of multiple optimal solutions, if they exist. For all these methods, I describe the algorithms used, along with a theoretical analysis of their running time and their statistical consistency. I also show results on biological and simulated data that demonstrate these methods’ effectiveness over a wide range of model conditions. In addition, I present the results of an experiment that compares various methods on trees simulated under both incomplete lineage sorting (ILS) as well as horizontal gene transfer (HGT).","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Pranjal Vachaspati, accepted the attached license on 2019-11-10 at 10:05.","The student, Pranjal Vachaspati, submitted this Dissertation for approval on 2019-11-10 at 10:09.","This Dissertation was approved for publication on 2019-12-01 at 10:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14537 on 2020-02-28 at 17:13:16","Made available in DSpace on 2020-03-02T21:58:07Z (GMT). No. of bitstreams: 3 VACHASPATI-DISSERTATION-2019.pdf: 4184201 bytes, checksum: 9e8727f177dc553edb24edc3eff93b72 (MD5) LICENSE.txt: 4215 bytes, checksum: 99c16004a270055ca4332c9ba0560981 (MD5) PROQUEST_LICENSE.txt: 4561 bytes, checksum: d72f0ab5308c2d622a3731ca22e4241b (MD5) Previous issue date: 2019-12-01"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large scale phylogenomic estimation"]}]}],"canonical_facts":{"dc:contributor":["Warnow, Tandy","Amato, Nancy","Chekuri, Chandra","Leebens-Mack, James"],"dc:creator":["Vachaspati, Pranjal"],"dc:date":["2020-03-02T21:58:07Z","2019-12-01","2019-12"],"dc:description":["Phylogenomic estimation - the science of calculating evolutionary trees from genomic data - is an important biological problem. As the amount of genomic data in biological datasets increases, new methods are needed to analyze this data. Cutting edge analyses may utilize genomes from tens of thousands of species. I present several methods for supertree and species tree estimation: ASTRID, FastRFS, SVDquest, and SIESTA. ASTRID can be used for both species tree and supertree estimation, and is designed to scale to very large datasets while maintaining a high level of accuracy. FastRFS is a supertree method that uses an exact constrained optimization algorithm to find accurate supertrees. SVDquest is a coalescent-aware species tree estimation method that estimates trees directly from sequences without using gene trees. Finally, SIESTA is a modification to the algorithms used by FastRFS, SVDquest, and other methods including ASTRAL that allows for the output and analysis of multiple optimal solutions, if they exist. For all these methods, I describe the algorithms used, along with a theoretical analysis of their running time and their statistical consistency. I also show results on biological and simulated data that demonstrate these methods’ effectiveness over a wide range of model conditions. In addition, I present the results of an experiment that compares various methods on trees simulated under both incomplete lineage sorting (ILS) as well as horizontal gene transfer (HGT).","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Pranjal Vachaspati, accepted the attached license on 2019-11-10 at 10:05.","The student, Pranjal Vachaspati, submitted this Dissertation for approval on 2019-11-10 at 10:09.","This Dissertation was approved for publication on 2019-12-01 at 10:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14537 on 2020-02-28 at 17:13:16","Made available in DSpace on 2020-03-02T21:58:07Z (GMT). No. of bitstreams: 3 VACHASPATI-DISSERTATION-2019.pdf: 4184201 bytes, checksum: 9e8727f177dc553edb24edc3eff93b72 (MD5) LICENSE.txt: 4215 bytes, checksum: 99c16004a270055ca4332c9ba0560981 (MD5) PROQUEST_LICENSE.txt: 4561 bytes, checksum: d72f0ab5308c2d622a3731ca22e4241b (MD5) Previous issue date: 2019-12-01"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106182"],"dc:language":["en"],"dc:rights":["Copyright 2019 Pranjal Vachaspati"],"dc:subject":["phylogenomics","species trees","computational biology","phylogenetics","evolution","incomplete lineage sorting","multispecies coalescent"],"dc:title":["Large scale phylogenomic estimation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}