University of Cambridge
Enabling the discovery of novel bioactives and enzymes from billions of protein sequences
Abstract
dc:description.abstractMicroorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Langer, Felix
- Advisor dc:contributor.advisor
-
- Finn, Robert
Subjects
dc:subject × 3Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.119801
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/386696