Back to results

University of Cambridge

Harnessing Deep Learning with Protein Language Models to Unveil Microbial Enzyme Function in Health and Disease

Abstract

dc:description.abstract

In microbial genomics, accurate annotations of the biological functions of enzymes are critical, as these proteins have important roles in catalysing essential biochemical reactions with high specificity and efficiency. Historically, functional annotation tools have relied on hidden Markov models (HMMs) that are built by aligning many amino acid sequences or using sequence homology tools like BLAST, which employs a pairwise alignment strategy between query and target sequences. Advancements in deep learning have significantly aided the functional annotation of proteins and comprehension of their diverse functions. Protein language models (pLMs), such as those used for structural prediction and other tasks, demonstrate remarkable capabilities in decoding the intricate amino acid language of proteins, which facilitates their functional annotation through a distinct approach compared to sequence-based alignment methods. In this body of work, I delve into the application and refinement of deep learning and classical machine learning techniques to unveil novel microbial enzymes with implications for health and disease. In the second chapter I combined pLMs with structural tools to create a new homology search system for annotating uncharacterized microbial proteins in IBD. I design several benchmarks for the evaluation of performing nearest neighbour searches to find protein homologs using the pLM embeddings. I demonstrate that the system can integrate protein structure when computing homology searches. My system uncovered potentially important enzymes in the IBD microbiome, including plsA, involved in plasmalogen biosynthesis. Additionally, I identified a remote homolog of McbB, a Pictet-Spengler enzyme, using structural search. McbB is hypothesised to synthesise ligands for GPR35, a receptor genetically linked to IBD. These discoveries, facilitated by deep learning tools, could offer new insights into IBD pathology and lead to novel therapeutic targets. In the third chapter I detail the development of CAZyLingua, the first tool that harnesses transfer learning from pLM embeddings to build a deep learning framework that facilitates the annotation of Carbohydrate Active Enzymes (CAZymes) in metagenomic datasets. I applied CAZyLingua to a paired mother/infant longitudinal dataset and revealed unannotated CAZymes linked to microbiome development during infancy. When applied to metagenomic datasets derived from patients affected by fibrosis-prone diseases such as Crohn’s disease and IgG4-related disease, CAZyLingua uncovered CAZymes associated with disease and healthy states. A CAZyme abundant in Crohn’s disease that CAZyLingua predicted to be a carbohydrate esterase was experimentally validated by demonstrating catalytic activity against acetylated manno-oligosaccharides. In the final chapter I describe connections between the gut and oral microbiomes to allergies. In particular, serine proteases (SPs) are emerging as potential allergens. This chapter uses deep learning and pre-trained pLMs to find allergenic SPs in metagenomic data. First, I develop a model to identify the key catalytic serine residue in serine hydrolases, showing how pLMs capture catalytic residues. Then, I build a deep learning framework to detect SP allergens across entire gene catalogues, using the structurally conserved catalytic triad to find potential homologs in gut and oral sites despite low sequence identity. The model predicts a putative SP allergen like a V8 protease, a known protease activated receptor-1 trigger. Notably, the model generalises beyond trained SP examples to predict a cysteine protease allergen resembling the Der f 1 dust mite allergen. This framework reveals allergens beyond those found by traditional sequence similarity, offering new targets for allergy research. This work represents a significant step towards redefining understanding of how microbial enzymes impact human health and disease. By integrating advanced deep learning, particularly pLMs, into existing functional annotation pipelines, previously hidden potential of enzymes within vast metagenomic datasets can be unlocked. This paves the way for the discovery of novel microbial enzyme functions associated with diverse health states, from IBD and chronic inflammatory diseases, to allergic responses and finally healthy physiological growth.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Thurimella, Kiran
Advisors dc:contributor.advisor
  • Owens, Roisin
  • Bacallado de Lara, Sergio

Subjects

dc:subject × 10

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.111559
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/372921

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Thurimella, Kiran. Harnessing Deep Learning with Protein Language Models to Unveil Microbial Enzyme Function in Health and Disease. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.111559