{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/374370"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/374370","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Machine learning in predictive toxicology: investigating developmental and reproductive toxicity with transfer learning","abstract":"The toxicity of a compound has always been a major concern in risk assessments of new products or drugs. In the field of predictive toxicology, with animal testing being phased out in some sectors, there is an urgent need for alternative methods for determining toxicity. An *in silico* method such as machine learning is one such popular choice given that the results can be obtained quickly with reasonable accuracy. Over the years, a number of machine learning models have been trained for various important human targets. However, machine learning models are limited by the quality of the data used to train them. In this work, the focus is on important toxicity endpoints that are being evaluated by next-generation risk assessments, including developmental and reproductive toxicity. Chapter 3 of this work investigates the use of Tanimoto similarity to determine the suitability of using transfer learning. It was found that when the predicted test accuracy (P) or the average similarity between datasets (S) is 70% or more, the machine learning model is likely to have high test accuracy when predicting on the test dataset. In Chapter 4, the creation of a new database for developmental and reproductive toxicity allows for newer machine learning models to be trained whose performances have been reported in this work. Models with about 68% accuracy for developmental toxicity and 80% for reproductive toxicity were trained. The suitability of transfer learning using the models for the two toxicity endpoints has also been investigated and several receptor bindings have been identified as possible mechanisms leading to developmental toxicity or reproductive toxicity.","abstract_html":"The toxicity of a compound has always been a major concern in risk assessments of new products or drugs. In the field of predictive toxicology, with animal testing being phased out in some sectors, there is an urgent need for alternative methods for determining toxicity. An *in silico* method such as machine learning is one such popular choice given that the results can be obtained quickly with reasonable accuracy. Over the years, a number of machine learning models have been trained for various important human targets. However, machine learning models are limited by the quality of the data used to train them. In this work, the focus is on important toxicity endpoints that are being evaluated by next-generation risk assessments, including developmental and reproductive toxicity. Chapter 3 of this work investigates the use of Tanimoto similarity to determine the suitability of using transfer learning. It was found that when the predicted test accuracy (P) or the average similarity between datasets (S) is 70% or more, the machine learning model is likely to have high test accuracy when predicting on the test dataset. In Chapter 4, the creation of a new database for developmental and reproductive toxicity allows for newer machine learning models to be trained whose performances have been reported in this work. Models with about 68% accuracy for developmental toxicity and 80% for reproductive toxicity were trained. The suitability of transfer learning using the models for the two toxicity endpoints has also been investigated and several receptor bindings have been identified as possible mechanisms leading to developmental toxicity or reproductive toxicity.","abstract_has_math":false,"creators":["Wang, Marcus"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Goodman, Jonathan"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-07-31","date_published":"2023-07-31","updated_at":"2026-07-22T22:24:13Z","subjects":["Computational chemistry","DART","Developmental and reproductive toxicology","Machine learning","Predictive toxicology","Transfer learning"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/2ea7eaf9-3991-4595-be83-9d731a6f21c8/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000284079275"],"render_values":[{"text":"0000-0002-8407-9275","href":"https://orcid.org/0000-0002-8407-9275","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.112456","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Goodman, Jonathan"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["We thank Unilever for their support and funding"]},{"key":"dc:creator","label":"Author","values":["Wang, Marcus"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000284079275"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2023-07-31"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/374370"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computational chemistry","DART","Developmental and reproductive toxicology","Machine learning","Predictive toxicology","Transfer learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/2ea7eaf9-3991-4595-be83-9d731a6f21c8/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.112456"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4ec6b4b-ae50-4e7e-bc34-f73808386389/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The toxicity of a compound has always been a major concern in risk assessments of new products or drugs. In the field of predictive toxicology, with animal testing being phased out in some sectors, there is an urgent need for alternative methods for determining toxicity. An *in silico* method such as machine learning is one such popular choice given that the results can be obtained quickly with reasonable accuracy. Over the years, a number of machine learning models have been trained for various important human targets. However, machine learning models are limited by the quality of the data used to train them. In this work, the focus is on important toxicity endpoints that are being evaluated by next-generation risk assessments, including developmental and reproductive toxicity. Chapter 3 of this work investigates the use of Tanimoto similarity to determine the suitability of using transfer learning. It was found that when the predicted test accuracy (P) or the average similarity between datasets (S) is 70% or more, the machine learning model is likely to have high test accuracy when predicting on the test dataset. In Chapter 4, the creation of a new database for developmental and reproductive toxicity allows for newer machine learning models to be trained whose performances have been reported in this work. Models with about 68% accuracy for developmental toxicity and 80% for reproductive toxicity were trained. The suitability of transfer learning using the models for the two toxicity endpoints has also been investigated and several receptor bindings have been identified as possible mechanisms leading to developmental toxicity or reproductive toxicity."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["e1680ade4b36600610d7328524b49755","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Machine learning in predictive toxicology: investigating developmental and reproductive toxicity with transfer learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Goodman, Jonathan"],"dc:contributor.sponsor":["We thank Unilever for their support and funding"],"dc:creator":["Wang, Marcus"],"dc:creator.authoridentifier":["0000000284079275"],"dc:date.issued":["2023-07-31"],"dc:description.abstract":["The toxicity of a compound has always been a major concern in risk assessments of new products or drugs. In the field of predictive toxicology, with animal testing being phased out in some sectors, there is an urgent need for alternative methods for determining toxicity. An *in silico* method such as machine learning is one such popular choice given that the results can be obtained quickly with reasonable accuracy. Over the years, a number of machine learning models have been trained for various important human targets. However, machine learning models are limited by the quality of the data used to train them. In this work, the focus is on important toxicity endpoints that are being evaluated by next-generation risk assessments, including developmental and reproductive toxicity. Chapter 3 of this work investigates the use of Tanimoto similarity to determine the suitability of using transfer learning. It was found that when the predicted test accuracy (P) or the average similarity between datasets (S) is 70% or more, the machine learning model is likely to have high test accuracy when predicting on the test dataset. In Chapter 4, the creation of a new database for developmental and reproductive toxicity allows for newer machine learning models to be trained whose performances have been reported in this work. Models with about 68% accuracy for developmental toxicity and 80% for reproductive toxicity were trained. The suitability of transfer learning using the models for the two toxicity endpoints has also been investigated and several receptor bindings have been identified as possible mechanisms leading to developmental toxicity or reproductive toxicity."],"dc:format.checksum.md5":["e1680ade4b36600610d7328524b49755","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.112456"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4ec6b4b-ae50-4e7e-bc34-f73808386389/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/374370"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/2ea7eaf9-3991-4595-be83-9d731a6f21c8/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Computational chemistry","DART","Developmental and reproductive toxicology","Machine learning","Predictive toxicology","Transfer learning"],"dc:title":["Machine learning in predictive toxicology: investigating developmental and reproductive toxicity with transfer learning"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:13Z"}