Back to results

University of Cambridge

Accurate Uncertainty Quantification and Explainable Artificial Intelligence in Machine Learning Models for Toxicological Risk Assessment

Abstract

dc:description.abstract

Consumer and environmental safety decisions can be supported by Quantitative Structure-Activity Relationship (QSAR) models – a key part of the Next Generation Risk Assessment strategy for animal-free safety. Machine learning methods are often employed to build QSAR models, but these “black box” functions still need to be validated robustly before being included in risk assessment strategies. Two key issues remain: uncertainty of the predictions and transparency of the model. The second chapter discusses mechanistically driven structural alerts for mitochondrial toxicity. Structural alerts are constructed using a maximum common substructure algorithm developed by Wedlake et al. (2020) and their mechanisms are verified by literature review. The alerts performed well on external validation when combined with existing structural alerts and can be further built upon in the future when more data is available. In chapter three, uncertainty quantification is studied by considering three different modelling methodologies (Bayesian bootstrapping, conformal prediction, and Bayesian neural networks) on a diverse dataset of 21 toxicologically relevant targets identified by Allen et al. (2022). Metrics to evaluate uncertainty quantification are defined and four interpretations of uncertainty are investigated. I show that being uncertain about a prediction does not necessarily imply higher error on average and epistemic uncertainty within a Bayesian neural network is correlated with the applicability domain of the model. Finally, in chapter four a Bayesian neural network is constructed using a large Ames mutagenicity dataset and evaluated on four different data splits based on source data, showing state of the art performance. Explanations for predictions are generated by two methods – RDKit SimilarityMaps based on molecular graph perturbation and SHAP applied to the Bayesian neural network. These explanations can reproduce existing literature structural alerts for covalent DNA binding developed by Enoch et al. (2010) and often where the model assigns a negative label it still correctly identifies the structural alert where applicable.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Gong, Chen
Advisor dc:contributor.advisor
  • Goodman, Jonathan

Subjects

dc:subject × 10

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.102202
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/358720

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Gong, Chen. Accurate Uncertainty Quantification and Explainable Artificial Intelligence in Machine Learning Models for Toxicological Risk Assessment. Doctoral thesis, University of Cambridge, 2023. https://doi.org/10.17863/CAM.102202