{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132486"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132486","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Integrating natural language with molecular structure","abstract":"The world faces unprecedented challenges in the coming decades across climate change, healthcare, and food security, each requiring innovative scientific solutions that are scalable, adaptable, and cost-effective. Further, we need to develop these solutions quickly. Broadly speaking, chemistry can provide molecular solutions to many of these problems: breakthrough drugs (e.g., kinase inhibitors), materials (e.g., organic photovoltaics), and chemical processes. The extremely large search spaces in which these solutions exist make AI tools critical for finding them. Of particular note, multimodal models combining natural, human language with molecular structure are poised to be a critical tool for discovering these solutions. In this dissertation, we will focus on enabling natural language processing, and particularly natural language itself, to serve as a tool for discovering and accelerating molecular solutions to critical global challenges. One of the first questions that probably comes to mind is why we would want to integrate natural language with molecules. Succinctly, combining these types of information has the possibility to accelerate scientific discovery. As motivating scenarios, imagine a future where a doctor can receive a novel, patient-specific drug necessary to treat an ailment just by writing a few sentences describing the patient’s symptoms (also taking into account their genotype, phenotype, and medical history). Or, imagine a scientist tackling challenging problems by designing a molecule satisfying desired functions (e.g., antimalarial or a photovoltaic) rather than its structure or low level properties (e.g., solubility). Controlling molecules and drug design in such a high-level manner has potential to be hugely impactful, but it requires a method of abstract description; luckily, humans have already developed one: natural language. As this research direction is a relatively new undertaking, we will focus on the development of several new tasks which are critical for its development. These include molecule captioning, text-based molecule generation, and inverse-synergistic drug structure design via in-context learning. Further, we will focus on three keys advantages of natural language in molecule design: functionality, abstraction, and composition. We explore both general approaches and specific applications to kinase inhibitor discovery, molecule property prediction, and drug synergy prediction. Finally, we conclude by proposing a modular chemical language model which is both synthesis- and function- aware. In particular, this model integrates lessons learned during earlier chapters by proposing a flexible, inference-time chemical vocabulary which can be adapted to a wide-variety of synthesis platforms.","abstract_html":"The world faces unprecedented challenges in the coming decades across climate change, healthcare, and food security, each requiring innovative scientific solutions that are scalable, adaptable, and cost-effective. Further, we need to develop these solutions quickly. Broadly speaking, chemistry can provide molecular solutions to many of these problems: breakthrough drugs (e.g., kinase inhibitors), materials (e.g., organic photovoltaics), and chemical processes. The extremely large search spaces in which these solutions exist make AI tools critical for finding them. Of particular note, multimodal models combining natural, human language with molecular structure are poised to be a critical tool for discovering these solutions. In this dissertation, we will focus on enabling natural language processing, and particularly natural language itself, to serve as a tool for discovering and accelerating molecular solutions to critical global challenges. One of the first questions that probably comes to mind is why we would want to integrate natural language with molecules. Succinctly, combining these types of information has the possibility to accelerate scientific discovery. As motivating scenarios, imagine a future where a doctor can receive a novel, patient-specific drug necessary to treat an ailment just by writing a few sentences describing the patient’s symptoms (also taking into account their genotype, phenotype, and medical history). Or, imagine a scientist tackling challenging problems by designing a molecule satisfying desired functions (e.g., antimalarial or a photovoltaic) rather than its structure or low level properties (e.g., solubility). Controlling molecules and drug design in such a high-level manner has potential to be hugely impactful, but it requires a method of abstract description; luckily, humans have already developed one: natural language. As this research direction is a relatively new undertaking, we will focus on the development of several new tasks which are critical for its development. These include molecule captioning, text-based molecule generation, and inverse-synergistic drug structure design via in-context learning. Further, we will focus on three keys advantages of natural language in molecule design: functionality, abstraction, and composition. We explore both general approaches and specific applications to kinase inhibitor discovery, molecule property prediction, and drug synergy prediction. Finally, we conclude by proposing a modular chemical language model which is both synthesis- and function- aware. In particular, this model integrates lessons learned during earlier chapters by proposing a flexible, inference-time chemical vocabulary which can be adapted to a wide-variety of synthesis platforms.","abstract_has_math":false,"creators":["Edwards, Carl"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Han, Jiawei","Zhai, ChengXiang","Burke, Martin D","Cho, Kyunghyun","Scalia, Gabriele"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["molecule-language multimodality","natural language processing","AI4Science","NLP4Science","molecule tokenization","molecule generation","LLM","chemical language model","drug synergy prediction","AI for scientific discovery","cross-modal learning","molecule-text alignment","chemical language models","molecule captioning","in-context molecular learning","text-to-molecule generation","text-guided molecule generation"],"languages":["en"],"rights":["Copyright 2025 Carl Edwards"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132486","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Han, Jiawei","Zhai, ChengXiang","Burke, Martin D","Cho, Kyunghyun","Scalia, Gabriele"]},{"key":"dc:creator","label":"Author","values":["Edwards, Carl"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-04"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["molecule-language multimodality","natural language processing","AI4Science","NLP4Science","molecule tokenization","molecule generation","LLM","chemical language model","drug synergy prediction","AI for scientific discovery","cross-modal learning","molecule-text alignment","chemical language models","molecule captioning","in-context molecular learning","text-to-molecule generation","text-guided molecule generation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Carl Edwards"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132486"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The world faces unprecedented challenges in the coming decades across climate change, healthcare, and food security, each requiring innovative scientific solutions that are scalable, adaptable, and cost-effective. Further, we need to develop these solutions quickly. Broadly speaking, chemistry can provide molecular solutions to many of these problems: breakthrough drugs (e.g., kinase inhibitors), materials (e.g., organic photovoltaics), and chemical processes. The extremely large search spaces in which these solutions exist make AI tools critical for finding them. Of particular note, multimodal models combining natural, human language with molecular structure are poised to be a critical tool for discovering these solutions. In this dissertation, we will focus on enabling natural language processing, and particularly natural language itself, to serve as a tool for discovering and accelerating molecular solutions to critical global challenges. One of the first questions that probably comes to mind is why we would want to integrate natural language with molecules. Succinctly, combining these types of information has the possibility to accelerate scientific discovery. As motivating scenarios, imagine a future where a doctor can receive a novel, patient-specific drug necessary to treat an ailment just by writing a few sentences describing the patient’s symptoms (also taking into account their genotype, phenotype, and medical history). Or, imagine a scientist tackling challenging problems by designing a molecule satisfying desired functions (e.g., antimalarial or a photovoltaic) rather than its structure or low level properties (e.g., solubility). Controlling molecules and drug design in such a high-level manner has potential to be hugely impactful, but it requires a method of abstract description; luckily, humans have already developed one: natural language. As this research direction is a relatively new undertaking, we will focus on the development of several new tasks which are critical for its development. These include molecule captioning, text-based molecule generation, and inverse-synergistic drug structure design via in-context learning. Further, we will focus on three keys advantages of natural language in molecule design: functionality, abstraction, and composition. We explore both general approaches and specific applications to kinase inhibitor discovery, molecule property prediction, and drug synergy prediction. Finally, we conclude by proposing a modular chemical language model which is both synthesis- and function- aware. In particular, this model integrates lessons learned during earlier chapters by proposing a flexible, inference-time chemical vocabulary which can be adapted to a wide-variety of synthesis platforms.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Carl Edwards, accepted the attached license on 2025-11-11 at 21:46.","The student, Carl Edwards, submitted this Dissertation for approval on 2025-12-03 at 22:05.","This Dissertation was approved for publication on 2025-12-04 at 08:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22860 on 2026-02-19 at 18:24:35"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Integrating natural language with molecular structure"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Han, Jiawei","Zhai, ChengXiang","Burke, Martin D","Cho, Kyunghyun","Scalia, Gabriele"],"dc:creator":["Edwards, Carl"],"dc:date":["2025-12","2025-12-04"],"dc:description":["The world faces unprecedented challenges in the coming decades across climate change, healthcare, and food security, each requiring innovative scientific solutions that are scalable, adaptable, and cost-effective. Further, we need to develop these solutions quickly. Broadly speaking, chemistry can provide molecular solutions to many of these problems: breakthrough drugs (e.g., kinase inhibitors), materials (e.g., organic photovoltaics), and chemical processes. The extremely large search spaces in which these solutions exist make AI tools critical for finding them. Of particular note, multimodal models combining natural, human language with molecular structure are poised to be a critical tool for discovering these solutions. In this dissertation, we will focus on enabling natural language processing, and particularly natural language itself, to serve as a tool for discovering and accelerating molecular solutions to critical global challenges. One of the first questions that probably comes to mind is why we would want to integrate natural language with molecules. Succinctly, combining these types of information has the possibility to accelerate scientific discovery. As motivating scenarios, imagine a future where a doctor can receive a novel, patient-specific drug necessary to treat an ailment just by writing a few sentences describing the patient’s symptoms (also taking into account their genotype, phenotype, and medical history). Or, imagine a scientist tackling challenging problems by designing a molecule satisfying desired functions (e.g., antimalarial or a photovoltaic) rather than its structure or low level properties (e.g., solubility). Controlling molecules and drug design in such a high-level manner has potential to be hugely impactful, but it requires a method of abstract description; luckily, humans have already developed one: natural language. As this research direction is a relatively new undertaking, we will focus on the development of several new tasks which are critical for its development. These include molecule captioning, text-based molecule generation, and inverse-synergistic drug structure design via in-context learning. Further, we will focus on three keys advantages of natural language in molecule design: functionality, abstraction, and composition. We explore both general approaches and specific applications to kinase inhibitor discovery, molecule property prediction, and drug synergy prediction. Finally, we conclude by proposing a modular chemical language model which is both synthesis- and function- aware. In particular, this model integrates lessons learned during earlier chapters by proposing a flexible, inference-time chemical vocabulary which can be adapted to a wide-variety of synthesis platforms.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Carl Edwards, accepted the attached license on 2025-11-11 at 21:46.","The student, Carl Edwards, submitted this Dissertation for approval on 2025-12-03 at 22:05.","This Dissertation was approved for publication on 2025-12-04 at 08:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22860 on 2026-02-19 at 18:24:35"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132486"],"dc:language":["en"],"dc:rights":["Copyright 2025 Carl Edwards"],"dc:subject":["molecule-language multimodality","natural language processing","AI4Science","NLP4Science","molecule tokenization","molecule generation","LLM","chemical language model","drug synergy prediction","AI for scientific discovery","cross-modal learning","molecule-text alignment","chemical language models","molecule captioning","in-context molecular learning","text-to-molecule generation","text-guided molecule generation"],"dc:title":["Integrating natural language with molecular structure"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}