University of Cambridge
Transfer learning in small-molecule drug discovery: From physics-based in-silico scores to realistic preclinical endpoints
Abstract
dc:description.abstractThis dissertation presents several projects in the general area of machine learning (ML) for drug discovery. Attempts to learn from molecular datasets are often hindered by the low availability of labeled datapoints. In the context of drug discovery, this limitation becomes especially severe due to the high cost of collecting experimental data. The common theme of this thesis is to improve model performance by transferring information from tasks where data is readily available to those where it is scarce. To this end, we pre-train or meta-train models on labels from public datasets or from physics-based simulations that are comparatively cheap to obtain. We find that transfer learning can greatly increase generalisation, even when downstream tasks are biologically unrelated to the original tasks. The dissertation includes one review chapter and four results chapters: 1. The first chapter reviews the application of ML to virtual screening of large chemical datasets which are too large to screen experimentally in the laboratory. 2. The second chapter describes initial experiments where we explore strategies to improve molecular representations from a variational autoencoder by training them to predict docking scores. Docking is a physics-inspired simulation of the binding mode between a small molecule and a protein. 3. The third chapter presents DOCKSTRING, a dataset of docking scores and a set of benchmark tasks for molecular ML based on docking simulations. DOCKSTRING data and tasks are useful in subsequent chapters. 4. The fourth chapter introduces graph neural processes for molecules. Neural processes are a family of models for meta-learning which try to emulate desirable properties of Gaussian processes with a neural architecture. In this chapter we evaluate their application to molecular datasets, benchmarking their performance in the DOCKSTRING tasks. We also explore strategies for fine-tuning parameters when confronted with novel tasks not previously seen during meta-training. 5. The fifth and final chapter explores the use of transfer learning to screen ultra-large chemical libraries in search of antibacterial compounds. We evaluated the amount of hits found when pre-training with different molecular datasets: docking scores, simple chemical properties computed in-silico or sparse bioactivity data from public datasets. After determining the optimal transfer-learning workflow, we purchased the top-ranking compounds from two commercial libraries and validated our predictions experimentally on the bacteria E. coli.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Garcia Ortegon, Miguel
- Advisors dc:contributor.advisor
-
- Bacallado, Sergio
- Bender, Andreas
- Rasmussen, Carl
Subjects
dc:subject × 3Rights
dc:rights- Licence
- Language dc:language
- eng
Identifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.115487
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/379366