{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/134196"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/134196","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Automated Synthesis Procedure Generation in Heterogeneous Catalysis via Fine-Tuned Language Models","abstract":"The exploration of catalytic materials and their synthesis routes traditionally demands extensive iterative experimentation and significant time investment. To overcome these constraints, we have developed an advanced extraction workflow integrating language models and multimodal processing techniques. Initially, textual data from over 9,000 scientific articles were analyzed to identify and extract detailed catalyst attributes such as chemical composition, structural motifs, morphology, crystal structure, size, shape, and support materials. Additionally, images and their associated captions were systematically captured from these publications, enriching the dataset through advanced vision- language processing methods. Subsequently, this structured information was refined through rigorous classification, synthesis query generation, and feasibility validation, resulting in a curated dataset comprising 1,632 high-quality catalyst synthesis procedures. Leveraging this dataset, we fine-tuned a large language model using parameter-efficient adaptation, significantly enhancing its capability to accurately predict detailed catalyst synthesis methods. Performance evaluation of our fine-tuned model revealed stable and effective convergence, demonstrating substantial improvements over baseline models with a ROUGE-1 score of 0.522, a ROUGE-L score of 0.290, and a BERTScore of 0.863. These results underscore the effectiveness of integrating multimodal data and validation methods, offering a powerful pathway to accelerate catalyst discovery, thereby reducing research timelines and resource demands.","abstract_html":"The exploration of catalytic materials and their synthesis routes traditionally demands extensive iterative experimentation and significant time investment. To overcome these constraints, we have developed an advanced extraction workflow integrating language models and multimodal processing techniques. Initially, textual data from over 9,000 scientific articles were analyzed to identify and extract detailed catalyst attributes such as chemical composition, structural motifs, morphology, crystal structure, size, shape, and support materials. Additionally, images and their associated captions were systematically captured from these publications, enriching the dataset through advanced vision- language processing methods. Subsequently, this structured information was refined through rigorous classification, synthesis query generation, and feasibility validation, resulting in a curated dataset comprising 1,632 high-quality catalyst synthesis procedures. Leveraging this dataset, we fine-tuned a large language model using parameter-efficient adaptation, significantly enhancing its capability to accurately predict detailed catalyst synthesis methods. Performance evaluation of our fine-tuned model revealed stable and effective convergence, demonstrating substantial improvements over baseline models with a ROUGE-1 score of 0.522, a ROUGE-L score of 0.290, and a BERTScore of 0.863. These results underscore the effectiveness of integrating multimodal data and validation methods, offering a powerful pathway to accelerate catalyst discovery, thereby reducing research timelines and resource demands.","abstract_has_math":false,"creators":["Diaz Aquino, Raul Bernardo"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Chemical Engineering","degree_department":"Chemical Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Xin, Hongliang"],"committee_members":["Bai, Xianming","Achenie, Luke E. K.","Deshmukh, Sanket A."],"year":2025,"date_issued":"2025-05-22","date_published":"2025-05-22","updated_at":"2026-07-22T22:19:35Z","subjects":["catalysis"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:43592"],"render_values":[{"text":"vt_gsexam:43592","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/134196","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Xin, Hongliang"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Bai, Xianming","Achenie, Luke E. K.","Deshmukh, Sanket A."]},{"key":"dc:contributor.department","label":"Department","values":["Chemical Engineering"]},{"key":"dc:creator","label":"Author","values":["Diaz Aquino, Raul Bernardo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-05-23T08:01:13Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-05-23T08:01:13Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-05-22"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Chemical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["catalysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:43592"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/134196"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The exploration of catalytic materials and their synthesis routes traditionally demands extensive iterative experimentation and significant time investment. To overcome these constraints, we have developed an advanced extraction workflow integrating language models and multimodal processing techniques. Initially, textual data from over 9,000 scientific articles were analyzed to identify and extract detailed catalyst attributes such as chemical composition, structural motifs, morphology, crystal structure, size, shape, and support materials. Additionally, images and their associated captions were systematically captured from these publications, enriching the dataset through advanced vision- language processing methods. Subsequently, this structured information was refined through rigorous classification, synthesis query generation, and feasibility validation, resulting in a curated dataset comprising 1,632 high-quality catalyst synthesis procedures. Leveraging this dataset, we fine-tuned a large language model using parameter-efficient adaptation, significantly enhancing its capability to accurately predict detailed catalyst synthesis methods. Performance evaluation of our fine-tuned model revealed stable and effective convergence, demonstrating substantial improvements over baseline models with a ROUGE-1 score of 0.522, a ROUGE-L score of 0.290, and a BERTScore of 0.863. These results underscore the effectiveness of integrating multimodal data and validation methods, offering a powerful pathway to accelerate catalyst discovery, thereby reducing research timelines and resource demands."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Digital technologies have significantly transformed the way we discover new catalysts and materials, starting a new era of scientific exploration. In our study, we introduce an innovative dual-model framework that leverages the strengths of diverse Large Language Models (LLMs) to accelerate research. We utilize an open source Large Language Model (LLM), a model highly adept at analyzing textual data, together with a Vision Language Model specifically designed for processing scientific images and their captions simultaneously. Through a comprehensive multi-step process, these models classify, extract, and structure valuable data on catalyst synthesis from over 9,000 peer-reviewed publications. The curated dataset is subsequently used to fine-tune another language model, through supervised training, enabling it to predict innovative synthesis pathways for catalysts and thus accelerate discovery. The collaborative operation between these models results in a robust extraction of synthesis mechanisms and offers new insights into catalytic processes. By integrating both textual and visual data, our approach not only accelerates the extraction of key scientific information but also increases the analysis of heterogeneous catalysis. This strategy aims to enhance catalyst synthesis efficiency, stimulate creative research directions, and contribute to significant advancements in materials science and chemical engineering. Overall, our findings shows how advanced digital tools can enhance catalyst research."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Automated Synthesis Procedure Generation in Heterogeneous Catalysis via Fine-Tuned Language Models"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Xin, Hongliang"],"dc:contributor.committeemember":["Bai, Xianming","Achenie, Luke E. K.","Deshmukh, Sanket A."],"dc:contributor.department":["Chemical Engineering"],"dc:creator":["Diaz Aquino, Raul Bernardo"],"dc:date.accessioned":["2025-05-23T08:01:13Z"],"dc:date.available":["2025-05-23T08:01:13Z"],"dc:date.issued":["2025-05-22"],"dc:description.abstract":["The exploration of catalytic materials and their synthesis routes traditionally demands extensive iterative experimentation and significant time investment. To overcome these constraints, we have developed an advanced extraction workflow integrating language models and multimodal processing techniques. Initially, textual data from over 9,000 scientific articles were analyzed to identify and extract detailed catalyst attributes such as chemical composition, structural motifs, morphology, crystal structure, size, shape, and support materials. Additionally, images and their associated captions were systematically captured from these publications, enriching the dataset through advanced vision- language processing methods. Subsequently, this structured information was refined through rigorous classification, synthesis query generation, and feasibility validation, resulting in a curated dataset comprising 1,632 high-quality catalyst synthesis procedures. Leveraging this dataset, we fine-tuned a large language model using parameter-efficient adaptation, significantly enhancing its capability to accurately predict detailed catalyst synthesis methods. Performance evaluation of our fine-tuned model revealed stable and effective convergence, demonstrating substantial improvements over baseline models with a ROUGE-1 score of 0.522, a ROUGE-L score of 0.290, and a BERTScore of 0.863. These results underscore the effectiveness of integrating multimodal data and validation methods, offering a powerful pathway to accelerate catalyst discovery, thereby reducing research timelines and resource demands."],"dc:description.abstractgeneral":["Digital technologies have significantly transformed the way we discover new catalysts and materials, starting a new era of scientific exploration. In our study, we introduce an innovative dual-model framework that leverages the strengths of diverse Large Language Models (LLMs) to accelerate research. We utilize an open source Large Language Model (LLM), a model highly adept at analyzing textual data, together with a Vision Language Model specifically designed for processing scientific images and their captions simultaneously. Through a comprehensive multi-step process, these models classify, extract, and structure valuable data on catalyst synthesis from over 9,000 peer-reviewed publications. The curated dataset is subsequently used to fine-tune another language model, through supervised training, enabling it to predict innovative synthesis pathways for catalysts and thus accelerate discovery. The collaborative operation between these models results in a robust extraction of synthesis mechanisms and offers new insights into catalytic processes. By integrating both textual and visual data, our approach not only accelerates the extraction of key scientific information but also increases the analysis of heterogeneous catalysis. This strategy aims to enhance catalyst synthesis efficiency, stimulate creative research directions, and contribute to significant advancements in materials science and chemical engineering. Overall, our findings shows how advanced digital tools can enhance catalyst research."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:43592"],"dc:identifier.uri":["https://hdl.handle.net/10919/134196"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["catalysis"],"dc:title":["Automated Synthesis Procedure Generation in Heterogeneous Catalysis via Fine-Tuned Language Models"],"dc:type":["Thesis"],"thesis:degree_discipline":["Chemical Engineering"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:35Z"}