Back to results

University of Cambridge

Transfer learning in small-molecule drug discovery: From physics-based in-silico scores to realistic preclinical endpoints

Abstract

dc:description.abstract

This dissertation presents several projects in the general area of machine learning (ML) for drug discovery. Attempts to learn from molecular datasets are often hindered by the low availability of labeled datapoints. In the context of drug discovery, this limitation becomes especially severe due to the high cost of collecting experimental data. The common theme of this thesis is to improve model performance by transferring information from tasks where data is readily available to those where it is scarce. To this end, we pre-train or meta-train models on labels from public datasets or from physics-based simulations that are comparatively cheap to obtain. We find that transfer learning can greatly increase generalisation, even when downstream tasks are biologically unrelated to the original tasks. The dissertation includes one review chapter and four results chapters: 1. The first chapter reviews the application of ML to virtual screening of large chemical datasets which are too large to screen experimentally in the laboratory. 2. The second chapter describes initial experiments where we explore strategies to improve molecular representations from a variational autoencoder by training them to predict docking scores. Docking is a physics-inspired simulation of the binding mode between a small molecule and a protein. 3. The third chapter presents DOCKSTRING, a dataset of docking scores and a set of benchmark tasks for molecular ML based on docking simulations. DOCKSTRING data and tasks are useful in subsequent chapters. 4. The fourth chapter introduces graph neural processes for molecules. Neural processes are a family of models for meta-learning which try to emulate desirable properties of Gaussian processes with a neural architecture. In this chapter we evaluate their application to molecular datasets, benchmarking their performance in the DOCKSTRING tasks. We also explore strategies for fine-tuning parameters when confronted with novel tasks not previously seen during meta-training. 5. The fifth and final chapter explores the use of transfer learning to screen ultra-large chemical libraries in search of antibacterial compounds. We evaluated the amount of hits found when pre-training with different molecular datasets: docking scores, simple chemical properties computed in-silico or sparse bioactivity data from public datasets. After determining the optimal transfer-learning workflow, we purchased the top-ranking compounds from two commercial libraries and validated our predictions experimentally on the bacteria E. coli.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Garcia Ortegon, Miguel
Advisors dc:contributor.advisor
  • Bacallado, Sergio
  • Bender, Andreas
  • Rasmussen, Carl

Subjects

dc:subject × 3

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.115487
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/379366

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Garcia Ortegon, Miguel. Transfer learning in small-molecule drug discovery: From physics-based in-silico scores to realistic preclinical endpoints. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.115487