Back to search

University of Toronto

Towards Data-driven Development of Advanced Drug Formulations Leveraging Machine Learning and Experimental Automation

Abstract

dc:description.abstract

The majority of new drugs fail during clinical trials, primarily due to ineffective performance, unacceptable side effects, and poor drug-like properties. Typically, many of these issues can be mitigated by developing fully optimized drug formulations. The development of these formulations usually involves extensive experimental work, starting with pre-formulation studies, followed by formulation design and optimization. Techniques, including molecular dynamics and design of experiment, have been employed to expedite this time-consuming and costly process. However, the limitations inherent in current methodologies and the growing complexity of advanced drug formulations highlight the need for innovative approaches. Over the past few decades, machine learning (ML) has emerged as a powerful tool in chemistry and materials science to significantly reduce the reliance on wet-lab experiments. More recently, its application in pharmaceutical sciences has shown promise in addressing similar challenges.This thesis aims to integrate ML into formulation development to accelerate the process of moving drugs from the bench to bedside. This work has included the development of ML models that assist in both pre-formulation and formulation optimization. For pre-formulation, an extensive dataset was collected from the existing literature. ML models were successfully developed using the collected dataset to predict the solubility of drugs, especially for those whose features align with the scope of the dataset. For formulation optimization, a dataset on drug release from long-acting injectable systems was compiled by combining data from formulations prepared and published by our group with data from the literature. Leveraging this dataset, an ML model was developed to predict fractional drug release and was interpreted to offer formulation insights. This thesis further investigated the potential of automation in generation of drug formulation data, using lipid-based nanoparticles as a case study. Specifically, an automated protocol was developed based on a liquid handling robot to generate a dataset in a high-throughput manner. ML models were then developed using this dataset to identify promising formulations, which exhibited in vivo performance comparable to commercial products. In summary, the findings from these studies highlight the potential of ML to accelerate the process of drug formulation. The integration of automation further facilitates this process by offering a cost-effective method for formulation screening and data curation.

Degree

thesis:*
Department dc:contributor.department
Pharmaceutical Sciences
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bao, Zeqing
Advisor dc:contributor.advisor
  • Allen, Christine

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • Attribution-NonCommercial-ShareAlike 4.0 International

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1807/140676
OAI identifier oai:identifier
oai:utoronto.scholaris.ca:1807/140676

Chain of custody

source
Harvested from
University of Toronto
Base URL
utoronto.scholaris.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Bao, Zeqing. Towards Data-driven Development of Advanced Drug Formulations Leveraging Machine Learning and Experimental Automation. 2024. http://hdl.handle.net/1807/140676