Back to results

University of Cambridge

Enabling the discovery of novel bioactives and enzymes from billions of protein sequences

Abstract

dc:description.abstract

Microorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Langer, Felix
Advisor dc:contributor.advisor
  • Finn, Robert

Subjects

dc:subject × 3

Rights

dc:rights

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.119801
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/386696

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Langer, Felix. Enabling the discovery of novel bioactives and enzymes from billions of protein sequences. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.119801