Back to results

University of Cambridge

Machine learning methods for detecting positive selection

Abstract

dc:description.abstract

Molecular evolutionary biology seeks to explain the diversity of life by uncovering the processes that shape genomes over time. A central aim within this field is to identify the genetic basis of adaptation, often manifesting as signatures of positive selection acting on protein-coding genes. Detecting such signals sheds light on evolutionary processes such as functional divergence, coevolutionary dynamics and phenotypic innovation. However, whilst the study of natural selection has been foundational in evolutionary theory, reliably identifying its genomic footprints remains a persistent challenge. Traditional methods for detecting interspecific positive selection are grounded in statistical, likelihood-based methods, typically employing codon substitution models. These approaches infer rates of nonsynonymous to synonymous substitutions (dN/dS) from nucleotide multiple sequence alignments (MSAs) of homologous, protein-coding genes, interpreted as a proxy for positive selection. These approaches have been instrumental in advancing our understanding of adaptive evolution, but they are not without their limitations. Statistical approaches are often computationally demanding, vulnerable to model misspecification, and underpowered when applied to realistic evolutionary scenarios. Models are forced to make simplifying assumptions to make the analysis tractable in the likelihood framework, which can contribute to both type I and type II error. Misalignments tend to further inflate false positive inference. As genomic datasets have grown in both scale and complexity, the shortcomings of classical inference methods have become increasingly apparent. This creates a pressing need for approaches that can harness large genomic datasets more flexibly, while capturing the complex dependencies inherent in sequence evolution. In parallel with these challenges, recent advances in machine learning have transformed diverse areas of biology, including molecular evolution. In particular, deep neural network architectures such as convolutional neural networks (CNNs) and transformer models have demonstrated an ability to extract meaningful patterns from raw biological data, without heavy reliance on human-curated features. Nevertheless, applying machine learning to molecular evolution presents its own challenges: the task requires large amounts of biologically realistic training data, architectures and resources equipped to handle such data, and rigorous evaluation frameworks that connect machine learning predictions to established evolutionary theory. There is very little or no evolutionary data for which we know the ground truth regarding how it has evolved; therefore, I rely on simulated data for training and evaluation. This thesis addresses these challenges by developing and evaluating machine learning methods for detecting positive selection in protein-coding genes. I first establish simulation frameworks that can be used to generate training data, and design benchmarking experiments to compare machine learning models against widely used statistical methods. I then develop CNNs trained to classify MSAs as evolving under positive selection or not, which were simulated from constrained phylogenies as a proof-of-principle. These models perform well when tested on data that resemble the training data, but struggle to generalise outside of the target domain. Nonetheless, these CNNs were successful in providing a faster, efficient and more accurate solution for positive selection detection compared to likelihood methods when trained for a specific phylogenetic scenario. This is useful for analyses such as genome-wide selection scans, where many MSAs are associated with the same phylogeny. To address the limitations of CNNs I then switch my focus to developing transformer models, which generalise more effectively across varied phylogenetic scenarios due to their global receptive field (ability to capture long-range dependencies across sequences) and permutation-invariant (insensitivity to the order of input elements) attention mechanism. I compare transformers directly with CNNs, showing comparable performance before extending to generalised models, capable of handling a wide range of phylogenetic scenarios. In addition to binary classification --- where whole genes are inferred to have evolved under positive selection or not --- I develop transformer models that predict sitewise dN/dS values for codon-specific inference of positive selection. To develop these models I evaluate a range of architectures and training strategies. These analyses demonstrate that machine learning approaches can achieve comparable or superior sensitivity to established likelihood-based methods. At the same time, I explore strategies to biologically interpret what these models learn. Exploiting the flexibility of the machine learning framework, I increase the realism of simulations by incorporating more realistic evolutionary processes that cannot typically be used in likelihood-based approaches. Finally, I apply generalised transformer models to empirical datasets. These analyses highlight their utility in cases where statistical methods lack power, while also revealing limitations that point to directions for future development. The contributions of this thesis are threefold. First, I provide an assessment of the limitations of current statistical methods, clarifying where they fall short in practice. Second, I develop new methodological tools that adapt machine learning architectures for evolutionary inference, with particular emphasis on CNNs and transformer models applied to MSAs. Third, I present an evaluation framework that bridges machine learning predictions with evolutionary theory, working towards methodological advances that are robust and biologically meaningful. Together, these contributions demonstrate how machine learning can extend the methodological toolkit of molecular evolution in ways that are difficult to achieve with traditional likelihood-based approaches. Machine learning models learn complex, non-linear patterns directly from data, capture signals of selection that are inaccessible to predefined parametric models, exploit the flexibility of simulation, and scale efficiently to large datasets. This work shows both the promise and the limitations of deep learning for detecting selection, and it points toward a future in which data-driven and theory-driven approaches are integrated to achieve faster, more flexible, and more accurate evolutionary inference.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • West, Charlotte
Advisor dc:contributor.advisor
  • Goldman, Nick

Subjects

dc:subject × 3

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.128461
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/400206

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

West, Charlotte. Machine learning methods for detecting positive selection. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.128461