{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113015"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113015","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large-scale methods for multiple sequence alignment and phylogeny estimation","abstract":"Bioinformatic analyses generally involve passing genetic sequence data through a pipeline of transformations. These prominently include multiple sequence alignment (MSA) and phylogeny estimation, which are both core challenges in computational biology. Their established mathematical formulations are usually difficult NP-hard problems that are attacked with careful and laborious heuristics. Consequently, the most accurate methods tend to be the most computationally expensive with increasing dataset size, and this limits the analysis to datasets of only modest volume. Larger datasets are becoming increasingly available and require qualitatively different approaches. In this thesis, we consider divide-and-conquer frameworks that employ slow-but-accurate base methods to efficiently solve the problem piecewise. The continued development of accurate and efficient large-scale methods will, it is hoped, facilitate more effective bioinformatic analysis of larger and larger datasets.","abstract_html":"Bioinformatic analyses generally involve passing genetic sequence data through a pipeline of transformations. These prominently include multiple sequence alignment (MSA) and phylogeny estimation, which are both core challenges in computational biology. Their established mathematical formulations are usually difficult NP-hard problems that are attacked with careful and laborious heuristics. Consequently, the most accurate methods tend to be the most computationally expensive with increasing dataset size, and this limits the analysis to datasets of only modest volume. Larger datasets are becoming increasingly available and require qualitatively different approaches. In this thesis, we consider divide-and-conquer frameworks that employ slow-but-accurate base methods to efficiently solve the problem piecewise. The continued development of accurate and efficient large-scale methods will, it is hoped, facilitate more effective bioinformatic analysis of larger and larger datasets.","abstract_has_math":false,"creators":["Smirnov, Vladimir A"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Warnow, Tandy","Peng, Jian","Forsyth, David","Treangen, Todd","Pop, Mihai"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:35Z","date_published":"2022-01-12T21:45:35Z","updated_at":"2026-07-22T22:24:52Z","subjects":["bioinformatics","computational biology","genetic sequences","multiple sequence alignment","phylogeny","big data"],"languages":["en"],"rights":["Copyright 2021 Vladimir Smirnov"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113015","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Warnow, Tandy","Peng, Jian","Forsyth, David","Treangen, Todd","Pop, Mihai"]},{"key":"dc:creator","label":"Author","values":["Smirnov, Vladimir A"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:35Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["bioinformatics","computational biology","genetic sequences","multiple sequence alignment","phylogeny","big data"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Vladimir Smirnov"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113015"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Bioinformatic analyses generally involve passing genetic sequence data through a pipeline of transformations. These prominently include multiple sequence alignment (MSA) and phylogeny estimation, which are both core challenges in computational biology. Their established mathematical formulations are usually difficult NP-hard problems that are attacked with careful and laborious heuristics. Consequently, the most accurate methods tend to be the most computationally expensive with increasing dataset size, and this limits the analysis to datasets of only modest volume. Larger datasets are becoming increasingly available and require qualitatively different approaches. In this thesis, we consider divide-and-conquer frameworks that employ slow-but-accurate base methods to efficiently solve the problem piecewise. The continued development of accurate and efficient large-scale methods will, it is hoped, facilitate more effective bioinformatic analysis of larger and larger datasets.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Vladimir Smirnov, accepted the attached license on 2021-07-10 at 13:12.","The student, Vladimir Smirnov, submitted this Dissertation for approval on 2021-07-10 at 13:29.","This Dissertation was approved for publication on 2021-07-12 at 11:24.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16848 on 2022-01-12 at 12:44:52","Made available in DSpace on 2022-01-12T21:45:35Z (GMT). No. of bitstreams: 4 SMIRNOV-DISSERTATION-2021.pdf: 3151076 bytes, checksum: 8f43a72e43c21069bf23d6be2181fb13 (MD5) Thesis.zip: 7108441 bytes, checksum: b5fadda52a96db5f88df8b5050fe2089 (MD5) LICENSE.txt: 4213 bytes, checksum: 672dc9068705b6d6ca94921a3a4ac2e4 (MD5) PROQUEST_LICENSE.txt: 4559 bytes, checksum: 49856eca1d784ecf3079f99d7e420b09 (MD5) Previous issue date: 2021-07-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large-scale methods for multiple sequence alignment and phylogeny estimation"]}]}],"canonical_facts":{"dc:contributor":["Warnow, Tandy","Peng, Jian","Forsyth, David","Treangen, Todd","Pop, Mihai"],"dc:creator":["Smirnov, Vladimir A"],"dc:date":["2022-01-12T21:45:35Z","2021-07-12","2021-08"],"dc:description":["Bioinformatic analyses generally involve passing genetic sequence data through a pipeline of transformations. These prominently include multiple sequence alignment (MSA) and phylogeny estimation, which are both core challenges in computational biology. Their established mathematical formulations are usually difficult NP-hard problems that are attacked with careful and laborious heuristics. Consequently, the most accurate methods tend to be the most computationally expensive with increasing dataset size, and this limits the analysis to datasets of only modest volume. Larger datasets are becoming increasingly available and require qualitatively different approaches. In this thesis, we consider divide-and-conquer frameworks that employ slow-but-accurate base methods to efficiently solve the problem piecewise. The continued development of accurate and efficient large-scale methods will, it is hoped, facilitate more effective bioinformatic analysis of larger and larger datasets.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Vladimir Smirnov, accepted the attached license on 2021-07-10 at 13:12.","The student, Vladimir Smirnov, submitted this Dissertation for approval on 2021-07-10 at 13:29.","This Dissertation was approved for publication on 2021-07-12 at 11:24.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16848 on 2022-01-12 at 12:44:52","Made available in DSpace on 2022-01-12T21:45:35Z (GMT). No. of bitstreams: 4 SMIRNOV-DISSERTATION-2021.pdf: 3151076 bytes, checksum: 8f43a72e43c21069bf23d6be2181fb13 (MD5) Thesis.zip: 7108441 bytes, checksum: b5fadda52a96db5f88df8b5050fe2089 (MD5) LICENSE.txt: 4213 bytes, checksum: 672dc9068705b6d6ca94921a3a4ac2e4 (MD5) PROQUEST_LICENSE.txt: 4559 bytes, checksum: 49856eca1d784ecf3079f99d7e420b09 (MD5) Previous issue date: 2021-07-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113015"],"dc:language":["en"],"dc:rights":["Copyright 2021 Vladimir Smirnov"],"dc:subject":["bioinformatics","computational biology","genetic sequences","multiple sequence alignment","phylogeny","big data"],"dc:title":["Large-scale methods for multiple sequence alignment and phylogeny estimation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}