{"id":{"repo_id":"texas","oai_identifier":"oai:repositories.lib.utexas.edu:2152/10540"},"canonical_url":"https://search.dev.ndltd.org/etd/texas/oai:repositories.lib.utexas.edu:2152/10540","repository":{"repo_id":"texas","name":"University of Texas","base_url":"https://repositories.lib.utexas.edu/server/oai/request"},"display":{"title":"Using a complex model of sequence evolution to evaluate and improve phylogenetic methods","abstract":"The performance of phylogenetic methods was evaluated by testing their success in recovering the true tree from computer simulated data. Data were generated on a variety of tree shapes under a complex model of sequence evolution based on the work of Aaron Halpern and William Bruno. Parameters of the model were estimated by maximum likelihood techniques applied to a phylogeny of 1,610 sequences of mammalian cytochrome b genes. These simulations represent a rigorous test of the robustness of phylogenetic methods because several of the simplifying assumptions made by inference methods are violated. Maximum likelihood methods assuming the general time reversible model of sequence evolution with rate heterogeneity proved to be quite robust, outperforming all other methods on small trees. Distance methods were significantly worse, even when implementing the same model of sequence evolution. On larger trees only distance methods and parsimony techniques were studied. In virtually all cases parsimony outperformed distance-based approaches. The use of simple distance corrections improved performance for four taxon trees and ultrametric sixteen taxon trees but decreased the performance of the neighbor joining method on a 228 taxon tree. Neighbor joining performed as well or better than searches under the minimum evolution criterion in all cases. In general, the results of these simulations agree with the conclusions of previous studies that phylogenetic methods perform well over a wide range of tree shapes, highly accurate phylogenies for large number of taxa can be obtained from moderate sequence lengths, and model-based distance corrections are much less robust than maximum likelihood implementations of the same model. Maximum likelihood under a common model of sequence evolution was found to be inconsistent on difficult tree shapes when the data were generated under the model described here. Analysis of the spectra of the generating model and the model of inference suggest slight alterations to the general time-reversible model that may improve its performance.","abstract_html":"The performance of phylogenetic methods was evaluated by testing their success in recovering the true tree from computer simulated data. Data were generated on a variety of tree shapes under a complex model of sequence evolution based on the work of Aaron Halpern and William Bruno. Parameters of the model were estimated by maximum likelihood techniques applied to a phylogeny of 1,610 sequences of mammalian cytochrome b genes. These simulations represent a rigorous test of the robustness of phylogenetic methods because several of the simplifying assumptions made by inference methods are violated. Maximum likelihood methods assuming the general time reversible model of sequence evolution with rate heterogeneity proved to be quite robust, outperforming all other methods on small trees. Distance methods were significantly worse, even when implementing the same model of sequence evolution. On larger trees only distance methods and parsimony techniques were studied. In virtually all cases parsimony outperformed distance-based approaches. The use of simple distance corrections improved performance for four taxon trees and ultrametric sixteen taxon trees but decreased the performance of the neighbor joining method on a 228 taxon tree. Neighbor joining performed as well or better than searches under the minimum evolution criterion in all cases. In general, the results of these simulations agree with the conclusions of previous studies that phylogenetic methods perform well over a wide range of tree shapes, highly accurate phylogenies for large number of taxa can be obtained from moderate sequence lengths, and model-based distance corrections are much less robust than maximum likelihood implementations of the same model. Maximum likelihood under a common model of sequence evolution was found to be inconsistent on difficult tree shapes when the data were generated under the model described here. Analysis of the spectra of the generating model and the model of inference suggest slight alterations to the general time-reversible model that may improve its performance.","abstract_has_math":false,"creators":["Holder, Mark Travis"],"institution":"The University of Texas at Austin","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Molecular Biology","degree_department":null,"school":null,"contributors":[],"advisors":["Hillis, David M., 1958-"],"committee_chairs":[],"committee_members":[],"year":2001,"date_issued":"2001-12","date_published":"2001-12","updated_at":"2026-07-24T05:01:08Z","subjects":["Phylogeny","Biology--Classification","Nucleotide sequence","Proteins--Analysis"],"languages":["eng"],"rights":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2152/10540","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hillis, David M., 1958-"]},{"key":"dc:creator","label":"Author","values":["Holder, Mark Travis"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2011-03-16T17:20:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2011-03-16T17:20:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2001-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Molecular Biology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Texas at Austin"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Phylogeny","Biology--Classification","Nucleotide sequence","Proteins--Analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/2152/10540"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["text"]},{"key":"dc:description.abstract","label":"Abstract","values":["The performance of phylogenetic methods was evaluated by testing their success in recovering the true tree from computer simulated data. Data were generated on a variety of tree shapes under a complex model of sequence evolution based on the work of Aaron Halpern and William Bruno. Parameters of the model were estimated by maximum likelihood techniques applied to a phylogeny of 1,610 sequences of mammalian cytochrome b genes. These simulations represent a rigorous test of the robustness of phylogenetic methods because several of the simplifying assumptions made by inference methods are violated. Maximum likelihood methods assuming the general time reversible model of sequence evolution with rate heterogeneity proved to be quite robust, outperforming all other methods on small trees. Distance methods were significantly worse, even when implementing the same model of sequence evolution. On larger trees only distance methods and parsimony techniques were studied. In virtually all cases parsimony outperformed distance-based approaches. The use of simple distance corrections improved performance for four taxon trees and ultrametric sixteen taxon trees but decreased the performance of the neighbor joining method on a 228 taxon tree. Neighbor joining performed as well or better than searches under the minimum evolution criterion in all cases. In general, the results of these simulations agree with the conclusions of previous studies that phylogenetic methods perform well over a wide range of tree shapes, highly accurate phylogenies for large number of taxa can be obtained from moderate sequence lengths, and model-based distance corrections are much less robust than maximum likelihood implementations of the same model. Maximum likelihood under a common model of sequence evolution was found to be inconsistent on difficult tree shapes when the data were generated under the model described here. Analysis of the spectra of the generating model and the model of inference suggest slight alterations to the general time-reversible model that may improve its performance."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["electronic"]},{"key":"dc:title","label":"Title","values":["Using a complex model of sequence evolution to evaluate and improve phylogenetic methods"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hillis, David M., 1958-"],"dc:creator":["Holder, Mark Travis"],"dc:date.accessioned":["2011-03-16T17:20:29Z"],"dc:date.available":["2011-03-16T17:20:29Z"],"dc:date.issued":["2001-12"],"dc:description":["text"],"dc:description.abstract":["The performance of phylogenetic methods was evaluated by testing their success in recovering the true tree from computer simulated data. Data were generated on a variety of tree shapes under a complex model of sequence evolution based on the work of Aaron Halpern and William Bruno. Parameters of the model were estimated by maximum likelihood techniques applied to a phylogeny of 1,610 sequences of mammalian cytochrome b genes. These simulations represent a rigorous test of the robustness of phylogenetic methods because several of the simplifying assumptions made by inference methods are violated. Maximum likelihood methods assuming the general time reversible model of sequence evolution with rate heterogeneity proved to be quite robust, outperforming all other methods on small trees. Distance methods were significantly worse, even when implementing the same model of sequence evolution. On larger trees only distance methods and parsimony techniques were studied. In virtually all cases parsimony outperformed distance-based approaches. The use of simple distance corrections improved performance for four taxon trees and ultrametric sixteen taxon trees but decreased the performance of the neighbor joining method on a 228 taxon tree. Neighbor joining performed as well or better than searches under the minimum evolution criterion in all cases. In general, the results of these simulations agree with the conclusions of previous studies that phylogenetic methods perform well over a wide range of tree shapes, highly accurate phylogenies for large number of taxa can be obtained from moderate sequence lengths, and model-based distance corrections are much less robust than maximum likelihood implementations of the same model. Maximum likelihood under a common model of sequence evolution was found to be inconsistent on difficult tree shapes when the data were generated under the model described here. Analysis of the spectra of the generating model and the model of inference suggest slight alterations to the general time-reversible model that may improve its performance."],"dc:format.medium":["electronic"],"dc:identifier.uri":["http://hdl.handle.net/2152/10540"],"dc:language.iso":["eng"],"dc:rights":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."],"dc:subject":["Phylogeny","Biology--Classification","Nucleotide sequence","Proteins--Analysis"],"dc:title":["Using a complex model of sequence evolution to evaluate and improve phylogenetic methods"],"thesis:degree_discipline":["Molecular Biology"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["The University of Texas at Austin"]},"updated_at":"2026-07-24T05:01:08Z"}