Universität Zurich
Development of an Improved Proteogenomics Solution to Identify Comprehensive Catalogues of Novel Small Proteins in Prokaryotes
Abstract
dc:descriptionProkaryotes, consisting of bacteria and archaea, are ubiquitous and diverse, playing key roles in highly relevant topics including human health and food production. Understanding the molecular mechanisms that underly their positive and negative effects in these areas is therefore an important topic of research and requires knowledge on the full sets of genes encoded in prokaryotic genomes, revealing all proteins these organisms can express to carry out biological functions. The small size and simple organization of prokaryotic genomes, in combination with advances in DNA sequencing technologies, has led to an exponential increase in prokaryotic organisms whose complete genome sequence is known. However, the identification of the genes encoded therein is still challenging, limiting the insights that can be gained. Genes are annotated using computational tools and ideally validated experimentally, but both approaches can miss genes, especially small ones. While small genes are more difficult to identify, the proteins they code for have been shown to be involved in important processes such as communication, antibiotics resistance and antimicrobial activity, and their identification has therefore become an active area of research. The most direct way to identify new protein-coding genes is to detect the corresponding protein using proteomics, which relies on mass spectrometry. This approach, called proteogenomics, requires a database which includes both the proteins that are known to be expressed by an organism and potential novel proteins. A tool to create such databases for prokaryotes by integrating different gene annotations and predictions called integrated proteogenomics databases (iPtgxDBs) has previously been developed by our team and was successfully used to identify new proteins in multiple prokaryotes. While successful, proteogenomics experiments can still only identify a fraction of the genes missed by genome annotation, necessitating innovations for a more comprehensive discovery of novel proteins. Furthermore, proteogenomics is still not widely used to improve genome annotations, in part because the bioinformatic analysis is challenging and not standardized. This thesis therefore focused on addressing both of these points. First, I explored an alternative proteomics approach that our collaborators developed specifically for the identification of small proteins, applying it to Pseudomonas stutzeri, a model system for denitrification. In contrast to a standard mass spectrometry workflow, this approach, called direct sequencing, does not rely on protease-based digestion but instead analyses whole proteins and their random decay products using an adapted mass spectrometry acquisition strategy. In combination with a tailor-made iPtgxDB that I generated based on the de novo assembled complete genome, this more than doubled the number of novel small proteins identified compared to a standard proteogenomics workflow. The enrichment of prolines in the sequences of these proteins demonstrates the benefit of direct sequencing for the identification of proteins that are difficult to identify with a standard trypsin digestion and three novel small proteins predicted to form an operon are one example of uncovered novelty that may provide new insights into P. stutzeri biology with further experimental analysis. Ribosome profiling detects actively translated mRNAs, providing a more sensitive prediction of new genes but with a higher potential for false positives compared to proteogenomics. As iPtgxDBs attempt to cover the full coding potential of an organism they are very large, which negatively impacts the protein identification rate. To improve the statistical power of proteogenomics searches, I therefore generated much smaller, custom iPtgxDBs for Sinorhizobium meliloti, a nitrogen fixing plant symbiont. The smaller database size was achieved by replacing the protein prediction source that has the lowest confidence and by far the largest number of predictions (an adapted six-frame translation) with the top candidates from a ribosome profiling analysis, for the first time established in this organism by our collaborators. The custom iPtgxDBs enabled the identification of additional annotated small proteins while the standard iPtgxDBs and ribosome profiling alone identified novel small proteins, a subset of which could be validated experimentally. The concept of such small, custom iPtgxDBs was further improved upon for the analysis of Synechocystis sp. PCC 6803, a cyanobacterium used as a model system for the study of photosynthesis. Our collaborators adapted a more sophisticated ribosome profiling approach for this organism that halts ribosomes at the translation initiation and termination sites, identifying exact gene boundaries. Based on this data, I again created a custom iPtgxDB and searched a very large collection of publicly available proteomics datasets against it. Since this required many searches, I automated large parts of the proteogenomics workflow and in the process integrated a new proteomics search engine and improved post-processing. Thanks to the extensiveness of the proteomics data, which covered multiple conditions, and the high quality of the ribosome profiling data, a substantial number of novel small proteins was identified both with a standard iPtgxDB and the custom iPtgxDB, many of which were unique to the respective database. This demonstrated the complementary benefits of both approaches and identified multiple novel genes of interest, including two putative antitoxins and a new family of proteins whose corresponding gene is contained within a longer, annotated gene. Experimental evidence showed that the pairs of longer and shorter proteins interact and may be functionally similar to a toxin-antitoxin system. This demonstrates the value of re-analyzing public proteomics data and custom iPtgxDBs, and confirms that even in a well-studied model organism such as Synechocystis sp. PCC 6803, novel insights based on ribosome profiling and proteogenomics can be gained. I then further extended the workflow for the analysis of phylogenomically closely related strains, applying it to six clinical reference strains of Mycobacterium tuberculosis, the deadliest bacterial infectious disease worldwide. A pan-genome analysis of the de novo assembled complete genomes revealed a large core genome in agreement with the known low genetic variability in M. tuberculosis. An automated pipeline I developed additionally enabled the identification and classification of PE and PPE genes. These genes are specific to M. tuberculosis and are often missed due to their repeat rich sequences but they may play important roles in virulence. The proteogenomics analysis relied for the first time on the widely available parallel accumulation–serial fragmentation (PASEF) mass spectrometry technology which increases the number of identified proteins. Besides integrating this technology, I further extended the proteogenomics workflow by implementing a pan-genomic analysis of the novel proteins, integrating improved peptide identification, and by incorporating a new approach to address the high rate of false positives among novel proteins identified by any proteogenomics study in a more data dependent manner. The identified novel proteins showed different levels of conservation between the six strains and included a putative antitoxin as well as a protein with structural similarity to a known antibiotic resistance protein. Finally, I created iPtgxDBs for a set of close to 47’000 complete prokaryotic genomes to make it easier for researchers to perform a proteogenomics study in their organism of interest. Based on the insights from earlier studies, I decided to create smaller custom iPtgxDBs, but since ribosome profiling data is only available for few prokaryotes, I evaluated the computational prediction of small protein candidates as an alternative. These custom iPtgxDBs, incorporating small protein predictions from the chosen tool, cover a wide taxonomic range, including species with non-standard genetic codes, and will be made publicly available. In conclusion, the explored experimental strategies all successfully enabled a more comprehensive identification of novel proteins and range from widely available approaches (PASEF) to more specialized methods (direct sequencing). The automated end-to-end proteogenomics workflow that was developed as part of this thesis supports all these strategies and implements important computational advancements such as better control of false positives among novel proteins. Its publication, combined with the public release of iPtgxDBs for a large number of species, will make proteogenomics based on iPtgxDBs, including the connected benefits, widely available to the research community and aid in the improvement of prokaryotic genome annotations, creating the basis for the discovery of novel proteins that can e.g., serve as drug targets, pharmaceuticals or biocontrol agents and consequently improve human health and food security.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Heiniger, Benjamin
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- Language dc:language
- eng