{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116180"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116180","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Machine learning for drug discovery and beyond","abstract":"The student, Wei Qian, accepted the attached license on 2022-07-01 at 14:44.","abstract_html":"The student, Wei Qian, accepted the attached license on 2022-07-01 at 14:44.","abstract_has_math":false,"creators":["Qian, Wei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian","El-Kebir, Mohammed","Han, Jiawei","Wiltschko, Alex"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["cheminformtics","drug discovery","computational biology","bioinformatics","machine learning","deep learning","artificial intelligence"],"languages":["en","eng"],"rights":["Copyright 2022 Wei Qian"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116180","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian","El-Kebir, Mohammed","Han, Jiawei","Wiltschko, Alex"]},{"key":"dc:creator","label":"Author","values":["Qian, Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-06"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["cheminformtics","drug discovery","computational biology","bioinformatics","machine learning","deep learning","artificial intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Wei Qian"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116180"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The student, Wei Qian, accepted the attached license on 2022-07-01 at 14:44.","The student, Wei Qian, submitted this Dissertation for approval on 2022-07-01 at 14:53.","This Dissertation was approved for publication on 2022-07-06 at 14:02.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","DSpace SAF Submission Ingestion Package generated from Vireo submission #18134 on 2022-11-15 at 17:38:05","The advent of digitized, large-scale, and high-throughput technologies has generated unprecedented data, presenting an excellent opportunity for today's drug discovery program to leverage machine learning (ML). By identifying relevant problems and suitable formulations in ML, we can translate these ever-increasing data into discovering better drugs and shorten the drug development cycles leading to cheaper drug and therapeutic options for previously incurable diseases. In this dissertation, we present four ML methods to tackle different challenges in today's drug discovery pipeline to quickly deliver more viable drug candidates for the clinical trial and eventually improve the quality of life for all humans. First, we introduce a batch equalization method that leverages style-transfer generative adversarial networks to mediate the batch effect commonly found in cellular images such that we can use them more effectively for high-throughput in vitro screening. Second, we describe an energy-inspired SE(3)-equivariant model to efficiently and accurately estimate the distribution of molecular conformations such that we can improve the accuracy for in silico structured-based screening. Third, we propose a 3D full-atom diffusion framework for target-aware molecule generation such that we can explore new chemistry beyond existing screening libraries and propose novel drug candidates for binding targets of challenging diseases. Fourth, we describe a reaction prediction algorithm that brings together rule-based systems (integer linear programming) and data-driven approaches (graph neural network) such that we can use efficiently synthesize drug candidates from the described screening pipeline or generative models. In the end, we use a graph neural network to model odorant molecules (instead of drugs) and find a universal odor space shared by many species. We hypothesize that the biology of metabolism drives such convergent evolution, and our ability to model these volatile organic compounds related to different metabolic processes could have great implications on how we understand animal olfaction and study human health. Put together, this dissertation shows the potential of machine learning to transform drug discovery and human health in the era of big data."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Machine learning for drug discovery and beyond"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian","El-Kebir, Mohammed","Han, Jiawei","Wiltschko, Alex"],"dc:creator":["Qian, Wei"],"dc:date":["2022-08","2022-07-06"],"dc:description":["The student, Wei Qian, accepted the attached license on 2022-07-01 at 14:44.","The student, Wei Qian, submitted this Dissertation for approval on 2022-07-01 at 14:53.","This Dissertation was approved for publication on 2022-07-06 at 14:02.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","DSpace SAF Submission Ingestion Package generated from Vireo submission #18134 on 2022-11-15 at 17:38:05","The advent of digitized, large-scale, and high-throughput technologies has generated unprecedented data, presenting an excellent opportunity for today's drug discovery program to leverage machine learning (ML). By identifying relevant problems and suitable formulations in ML, we can translate these ever-increasing data into discovering better drugs and shorten the drug development cycles leading to cheaper drug and therapeutic options for previously incurable diseases. In this dissertation, we present four ML methods to tackle different challenges in today's drug discovery pipeline to quickly deliver more viable drug candidates for the clinical trial and eventually improve the quality of life for all humans. First, we introduce a batch equalization method that leverages style-transfer generative adversarial networks to mediate the batch effect commonly found in cellular images such that we can use them more effectively for high-throughput in vitro screening. Second, we describe an energy-inspired SE(3)-equivariant model to efficiently and accurately estimate the distribution of molecular conformations such that we can improve the accuracy for in silico structured-based screening. Third, we propose a 3D full-atom diffusion framework for target-aware molecule generation such that we can explore new chemistry beyond existing screening libraries and propose novel drug candidates for binding targets of challenging diseases. Fourth, we describe a reaction prediction algorithm that brings together rule-based systems (integer linear programming) and data-driven approaches (graph neural network) such that we can use efficiently synthesize drug candidates from the described screening pipeline or generative models. In the end, we use a graph neural network to model odorant molecules (instead of drugs) and find a universal odor space shared by many species. We hypothesize that the biology of metabolism drives such convergent evolution, and our ability to model these volatile organic compounds related to different metabolic processes could have great implications on how we understand animal olfaction and study human health. Put together, this dissertation shows the potential of machine learning to transform drug discovery and human health in the era of big data."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116180"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Wei Qian"],"dc:subject":["cheminformtics","drug discovery","computational biology","bioinformatics","machine learning","deep learning","artificial intelligence"],"dc:title":["Machine learning for drug discovery and beyond"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}