University of Cambridge
A gene-space de Bruijn graph framework for improved genotyping of antimicrobial resistance genes
Abstract
dc:description.abstractIdentifying antimicrobial resistance (AMR) in bacteria is critical for surveillance and clinical diagnosis. Conventional antimicrobial susceptibility testing involves culturing isolates and measuring growth across antimicrobial concentration gradients. However, this approach is slow and may yield results that differ between in vitro and in vivo conditions. As a result, genomics has become an important approach, enabling the detection of AMR-associated genes and single nucleotide polymorphisms. This typically involves assembling sequencing reads and annotating the assemblies for known AMR determinants. Long-read sequencing technologies have greatly improved the completeness of bacterial genome assemblies, allowing for more accurate detection of AMR. However, even with long-reads, assembly remains imperfect and this can result in the under-detection of genuinely present AMR genes and inaccurate predictions of antimicrobial susceptibility. Moreover, multi-copy AMR genes can have dosage effects, whereby the number of copies of a particular AMR gene can increase the concentration of antimicrobial needed to treat an infection. These types of genes often occur in repetitive regions of the genome that are prone to misassembly. This thesis focuses on the development and application of a novel de Bruijn graph–based data structure, the gene-space de Bruijn graph (gene DBG), for AMR gene detection. In this graph, the k-mer alphabet is the set of genes in the pan-genome of the bacterial species under analysis. This approach uses an existing tool, Pandora, to detect the genes on long-read sequences, then constructs and corrects gene DBGs where the nodes represent k adjacent genes, and edges connect nodes that occur consecutively in any read. This enables effective separation and clustering of reads corresponding to distinct genomic copies of multi-copy AMR genes based on their paths through the graph. The reads within each cluster are then used to reconstruct the nucleotide sequences for each AMR gene genomic copy and to estimate their cellular copy number to account multi-copy plasmids. This approach is implemented in a tool called Amira. This thesis first presents an in-depth evaluation of the gene detection accuracy of Pandora and the development of a workflow to automate the construction of high-quality pan-genome reference graphs (panRGs) for panels of clinically relevant bacterial species. The methodology for constructing and correcting gene DBGs is then introduced and systematically evaluated. Building on this, an approach is developed to identify and exploit the paths traversed by reads through AMR gene–containing nodes in the gene DBG to cluster the reads containing AMR genes. This enables direct estimation of AMR gene genomic copy number, cellular copy number and nucleotide sequence directly from long-read sequencing data. The accuracy of Amira is systematically compared to alternative AMR gene detection approaches using simulated and manually curated single-isolate reference datasets. This demonstrates that Amira achieves higher accuracy in resolving AMR gene presence and genomic copy number than assembly-based methods. A large-scale analysis of thousands of bacterial genomes further shows that Amira is more sensitive in detecting AMR gene presence than a leading genome assembler, which systematically misassembles AMR genes and underestimates their frequencies. Finally, Amira is shown to be applicable to AMR gene detection in complex communities and performs robustly on both a simulated mixed-strain dataset and a mock metagenomic community. This work introduces the gene DBG, a novel data structure that substantially improves the accuracy of AMR gene detection over existing methods. We anticipate that both this framework and its software implementation will be of broad interest to the scientific community and will support more accurate AMR gene detection in future epidemiological studies.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Anderson, Daniel
- Advisor dc:contributor.advisor
-
- Iqbal, Zamin
Subjects
dc:subject × 6Rights
dc:rights- Licence
- Language dc:language
- eng
Identifiers
dc:identifier.*- Author Identifier
- 0000-0003-4422-9520
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/398737