Back to results

Universität Tübingen

Biologically meaningful classification of protein sequences - a bioinformatic approach

Abstract

Life without proteins is hardly imaginable. Proteins are essential to most structural components and metabolic processes within cells and replication of genetic material would not be possible, were they missing. The Genome of each organism contains information about all proteins that organism is capable of synthesizing. As proteins are such a central component of life, it is essential to gain a greater unterstanding of the various proteins and their interaction partners, prior to being able to understand Organisms at a molecular resolution. Experimental characterization of all proteins in all organisms is unfeasable due to time and financial constraints. However, it is frequently possible to glean knowledge for a large number of proteins in each new genome by transferring information from close sequence relatives wich have been characterized. The idea being, that proteins similar at the sequence level will most likey also have retained a similar structure and function. Some of the experimentally determined characteristics of one protein can therefore be transferred to all related proteins, depending on the degree of relatedness. Protein classification deals with determining the degree to which proteins are related and which functional and structural characteristics are conserved. In this work I describe the basics of protein classification: sequence similarity searches, sequence alignment and phylogenetic inference. Various methods are described and the advantages and disadvantages of one approach over the other mentioned. In addition, the most frequent protein classification problems and ways to circumvent these are presented. PhyloGenie and CLANS describe two different approaches to protein classification. Phylogenie focuses on the analysis of the set of all trees derived from the proteome of an organism: the phylome. To compare the performance of phylogenie to alternative methods, we repeated the analysis of two datasets searching for: 1) the amount of lateral gene transfer between Thermoplasma and Sulfolobus (Ruepp et al. 2000) and 2) genes supporting the hypothesis of an actinopterygian specific genome duplication (Taylor et al. 2003). Our analysis of the Thermoplasma acidophilum dataset pointed to large numbers of genes having been transferred between Thermoplasma and distantly related archaebacteria of the genus Sulfolobus. Comparison with other methods of detecting lateral gene transfer showed PhyloGenie to provide the best sensitivity to specificity quotient of the tested methods. Using Phylogenie in a comparative genomics analysis of the incomplete Dario rerio genome, we were able to double the number of orthologous genes supporting the actinopterygian specific genome duplication hypothesis. In contrast to PhyloGenie, which works in a mostly organism-specific manner, CLANS is used to analyze protein families. Protein families are used to describe the set of sequences descendant from an ancestral protein, some of which may have greatly changed over time. Larger families may contain orthologous and paralogous subgroups and encompass many thousands of sequences, rendering phylogenetic approaches computationally prohibitive and difficult to analyze. CLANS relies on graphical representation of all pairwise sequence similarities. This permits analysis of much larger datasets and is less sensitive to many of the problems traditional phylogenetic methods face. Application of CLANS to the group of AAA-ATPases enabled us to describe this family in an objective manner for the first time. Previous analyses differed in number and types of sequences used, so that enumeration and classification of all AAA-ATPases in the NCBI nonredundant protein database was a primary goal. The results generated were biologically plausible and surprising insights, such as the apparent homology of N-domains of distantly related AAA-ATPases, could be corroborated by additional tests. Due to it's ability to rapidly analyze large numbers of unaligned sequences, CLANS became the basis for a number of further analyses. Published examples include a description of the TAA43 protein (Santos et al. 2004), the Wipi-1-alpha beta-propeller (Proikas-Czesanne et al. 2004) as well as a correction of the structure of the AbrB transcription factor (Coles et al. in press).

Author and committee

dc:creator, dc:contributor.*
Author
  • Frickey, Tancred Gilles

Identifiers

dc:identifier.*
Identifier
hdl:10900/48804

Chain of custody

source
Harvested from
Universität Tübingen
Base URL
publikationen.uni-tuebingen.de/oai/request
Last updated
2026-08-21
Source record
OAI-PMH GetRecord
related terms
citation

Frickey, Tancred Gilles. Biologically meaningful classification of protein sequences - a bioinformatic approach. 2005.