{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105567"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105567","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical estimation problems in phylogenomics and applications in microbial ecology","abstract":"With the growing awareness of the potential for microbial communities to play a role in human health, environmental remediation and other important processes, the challenge of understanding such a complex population through the lens of high-throughput sequencing output has risen to the fore. For a de novo sequenced community, the first step to understanding the population involves comparing the sequences to a reference database in some form. In this dissertation, we consider some challenges and benefits of organizing the reference data according to evolution, with orthologous genes grouped together and stored as a multiple sequence alignment and phylogenetic tree. First we consider the related problem of estimating the population-level phylogeny of a group of species based on the alignments and phylogenies of several individual genes. Under one common model, species tree estimation is provably statistically consistent by several different methods, but those proofs rely on two separate and potentially shaky assumptions: that every species appears in the data for every gene (i.e., there is no missing data), and that since gene tree estimation is itself consistent, the gene trees used to compute the population-level tree are correct. Second, we explore some novel ways to use a Bayesian MCMC algorithm for jointly estimating alignment and phylogeny. The result is increased accuracy for large alignments, where the MCMC method alone would not be tractable. In the process, we identify a peculiar property of this Bayesian algorithm: it performs much differently on simulated sequences than on sequences from biological alignment benchmarks. No other alignment method tested showed the same divergence. Finally, we present two different practical applications a reference database containing an alignment and tree for a group of gene families in the context of microbial ecology. The first is an algorithm that uses the tree and alignment to construct an ensemble of profile hidden Markov models that improves remote homology detection. The second is a data visualization technique that generates an image of the community with a high density of data, but one that makes it naturally easy to compare many different samples at a time, potentially uncovering otherwise elusive patterns in the data.","abstract_html":"With the growing awareness of the potential for microbial communities to play a role in human health, environmental remediation and other important processes, the challenge of understanding such a complex population through the lens of high-throughput sequencing output has risen to the fore. For a de novo sequenced community, the first step to understanding the population involves comparing the sequences to a reference database in some form. In this dissertation, we consider some challenges and benefits of organizing the reference data according to evolution, with orthologous genes grouped together and stored as a multiple sequence alignment and phylogenetic tree. First we consider the related problem of estimating the population-level phylogeny of a group of species based on the alignments and phylogenies of several individual genes. Under one common model, species tree estimation is provably statistically consistent by several different methods, but those proofs rely on two separate and potentially shaky assumptions: that every species appears in the data for every gene (i.e., there is no missing data), and that since gene tree estimation is itself consistent, the gene trees used to compute the population-level tree are correct. Second, we explore some novel ways to use a Bayesian MCMC algorithm for jointly estimating alignment and phylogeny. The result is increased accuracy for large alignments, where the MCMC method alone would not be tractable. In the process, we identify a peculiar property of this Bayesian algorithm: it performs much differently on simulated sequences than on sequences from biological alignment benchmarks. No other alignment method tested showed the same divergence. Finally, we present two different practical applications a reference database containing an alignment and tree for a group of gene families in the context of microbial ecology. The first is an algorithm that uses the tree and alignment to construct an ensemble of profile hidden Markov models that improves remote homology detection. The second is a data visualization technique that generates an image of the community with a high density of data, but one that makes it naturally easy to compare many different samples at a time, potentially uncovering otherwise elusive patterns in the data.","abstract_has_math":false,"creators":["Nute, Michael Gordon"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Warnow, Tandy J","Gropp, William","Zhao, Dave","Stumpf, Rebecca","Chen, Yuguo","Pop, Mihai"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:28:48Z","date_published":"2019-11-26T20:28:48Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Phylogenetics","Mutliple Sequence Alignment"],"languages":["en"],"rights":["Copyright 2019 Michael Nute"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105567","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Warnow, Tandy J","Gropp, William","Zhao, Dave","Stumpf, Rebecca","Chen, Yuguo","Pop, Mihai"]},{"key":"dc:creator","label":"Author","values":["Nute, Michael Gordon"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:28:48Z","2019-04-19","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Phylogenetics","Mutliple Sequence Alignment"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Michael Nute"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105567"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["With the growing awareness of the potential for microbial communities to play a role in human health, environmental remediation and other important processes, the challenge of understanding such a complex population through the lens of high-throughput sequencing output has risen to the fore. For a de novo sequenced community, the first step to understanding the population involves comparing the sequences to a reference database in some form. In this dissertation, we consider some challenges and benefits of organizing the reference data according to evolution, with orthologous genes grouped together and stored as a multiple sequence alignment and phylogenetic tree. First we consider the related problem of estimating the population-level phylogeny of a group of species based on the alignments and phylogenies of several individual genes. Under one common model, species tree estimation is provably statistically consistent by several different methods, but those proofs rely on two separate and potentially shaky assumptions: that every species appears in the data for every gene (i.e., there is no missing data), and that since gene tree estimation is itself consistent, the gene trees used to compute the population-level tree are correct. Second, we explore some novel ways to use a Bayesian MCMC algorithm for jointly estimating alignment and phylogeny. The result is increased accuracy for large alignments, where the MCMC method alone would not be tractable. In the process, we identify a peculiar property of this Bayesian algorithm: it performs much differently on simulated sequences than on sequences from biological alignment benchmarks. No other alignment method tested showed the same divergence. Finally, we present two different practical applications a reference database containing an alignment and tree for a group of gene families in the context of microbial ecology. The first is an algorithm that uses the tree and alignment to construct an ensemble of profile hidden Markov models that improves remote homology detection. The second is a data visualization technique that generates an image of the community with a high density of data, but one that makes it naturally easy to compare many different samples at a time, potentially uncovering otherwise elusive patterns in the data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Michael Nute, accepted the attached license on 2019-04-18 at 16:57.","The student, Michael Nute, submitted this Dissertation for approval on 2019-04-18 at 17:05.","This Dissertation was approved for publication on 2019-04-19 at 13:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13759 on 2019-11-26 at 12:47:54","Made available in DSpace on 2019-11-26T20:28:48Z (GMT). No. of bitstreams: 5 NUTE-DISSERTATION-2019.pdf: 5060925 bytes, checksum: 2ed911d6cda212e34969254a7d9e1110 (MD5) PICAN-PI_supplement_Primate_Vaginal_all.zip: 48833518 bytes, checksum: 14ec570c381a30b947b8751c9a9ea497 (MD5) PICAN_PI_supplement_IBD_animations_all.zip: 27835138 bytes, checksum: b6c4589adc9c1ad7018d244b326fa095 (MD5) LICENSE.txt: 4209 bytes, checksum: 90e3c5679b1278a584918d4ab01560b9 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 2251de5deb55a61b414e9452532da67e (MD5) Previous issue date: 2019-04-19"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Statistical estimation problems in phylogenomics and applications in microbial ecology"]}]}],"canonical_facts":{"dc:contributor":["Warnow, Tandy J","Gropp, William","Zhao, Dave","Stumpf, Rebecca","Chen, Yuguo","Pop, Mihai"],"dc:creator":["Nute, Michael Gordon"],"dc:date":["2019-11-26T20:28:48Z","2019-04-19","2019-08"],"dc:description":["With the growing awareness of the potential for microbial communities to play a role in human health, environmental remediation and other important processes, the challenge of understanding such a complex population through the lens of high-throughput sequencing output has risen to the fore. For a de novo sequenced community, the first step to understanding the population involves comparing the sequences to a reference database in some form. In this dissertation, we consider some challenges and benefits of organizing the reference data according to evolution, with orthologous genes grouped together and stored as a multiple sequence alignment and phylogenetic tree. First we consider the related problem of estimating the population-level phylogeny of a group of species based on the alignments and phylogenies of several individual genes. Under one common model, species tree estimation is provably statistically consistent by several different methods, but those proofs rely on two separate and potentially shaky assumptions: that every species appears in the data for every gene (i.e., there is no missing data), and that since gene tree estimation is itself consistent, the gene trees used to compute the population-level tree are correct. Second, we explore some novel ways to use a Bayesian MCMC algorithm for jointly estimating alignment and phylogeny. The result is increased accuracy for large alignments, where the MCMC method alone would not be tractable. In the process, we identify a peculiar property of this Bayesian algorithm: it performs much differently on simulated sequences than on sequences from biological alignment benchmarks. No other alignment method tested showed the same divergence. Finally, we present two different practical applications a reference database containing an alignment and tree for a group of gene families in the context of microbial ecology. The first is an algorithm that uses the tree and alignment to construct an ensemble of profile hidden Markov models that improves remote homology detection. The second is a data visualization technique that generates an image of the community with a high density of data, but one that makes it naturally easy to compare many different samples at a time, potentially uncovering otherwise elusive patterns in the data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Michael Nute, accepted the attached license on 2019-04-18 at 16:57.","The student, Michael Nute, submitted this Dissertation for approval on 2019-04-18 at 17:05.","This Dissertation was approved for publication on 2019-04-19 at 13:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13759 on 2019-11-26 at 12:47:54","Made available in DSpace on 2019-11-26T20:28:48Z (GMT). No. of bitstreams: 5 NUTE-DISSERTATION-2019.pdf: 5060925 bytes, checksum: 2ed911d6cda212e34969254a7d9e1110 (MD5) PICAN-PI_supplement_Primate_Vaginal_all.zip: 48833518 bytes, checksum: 14ec570c381a30b947b8751c9a9ea497 (MD5) PICAN_PI_supplement_IBD_animations_all.zip: 27835138 bytes, checksum: b6c4589adc9c1ad7018d244b326fa095 (MD5) LICENSE.txt: 4209 bytes, checksum: 90e3c5679b1278a584918d4ab01560b9 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 2251de5deb55a61b414e9452532da67e (MD5) Previous issue date: 2019-04-19"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105567"],"dc:language":["en"],"dc:rights":["Copyright 2019 Michael Nute"],"dc:subject":["Phylogenetics","Mutliple Sequence Alignment"],"dc:title":["Statistical estimation problems in phylogenomics and applications in microbial ecology"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}