{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129195"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129195","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"AI4Scientist: Accelerating and democratizing scientific research lifecycle","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Wang, Qingyun"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Han, Jiawei","Hakkani-Tur, Dilek","Zhao, Han","Neubig, Graham","Hope, Tom"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-15","date_published":"2025-04-15","updated_at":"2026-07-22T22:25:04Z","subjects":["Natrual language generation","Scientific information extraction","Scientific knowledge reasoning","Scientific knowledge dissemination","AI4Scientist"],"languages":["en","eng"],"rights":["Copyright 2025 Qingyun Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129195","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Han, Jiawei","Hakkani-Tur, Dilek","Zhao, Han","Neubig, Graham","Hope, Tom"]},{"key":"dc:creator","label":"Author","values":["Wang, Qingyun"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-15","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natrual language generation","Scientific information extraction","Scientific knowledge reasoning","Scientific knowledge dissemination","AI4Scientist"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Qingyun Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129195"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Qingyun Wang, accepted the attached license on 2025-04-11 at 19:28.","The student, Qingyun Wang, submitted this Dissertation for approval on 2025-04-11 at 19:28.","This Dissertation was approved for publication on 2025-04-15 at 09:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21750 on 2025-10-19 at 18:09:19","Millions of scientific papers are published annually, resulting in an information overload. Beyond this, scientific papers, known as \"Sleeping beauties\", sometimes remain largely unnoticed for long periods before suddenly attracting great attention. Moreover, the process of discovering new scientific hypotheses has remained slow, expensive, and highly specialist-dependent, due to the increasingly complex experiments. There exists a pressing need to help scientists digest and evaluate relevant papers, facilitating scientific discovery. However, computer-human collaboration in scientific hypothesis discovery is still exploratory and lacks a unified framework for analyzing the relevant tasks. This thesis tackles the problem of automating scientific literature understanding and scientific discovery by proposing an AI4Scientist. Unlike AI4Science, which creates AI-driven solutions for scientific challenges, we focus on AI4Scientists to empower scientists with AI tools to enhance their research lifecycle. The recent advancements in large language models (LLMs) raise the prospect that they may be able to solve those problems. Despite their impressive progress, LLMs often fail to effectively incorporate domain-specific knowledge and support their generation with enough evidence. Additionally, the expert-curated databases represent only a small fraction of knowledge in the entire domain, due to high annotation cost. To address this issue and lower the entry barrier for interdisciplinary collaboration, we develop AI tools to accelerate the entire research lifecycle for scientists, from knowledge acquisition, hypothesis generation, multimedia procedure planning for experiment design, experiment execution, conduction to writing, and evaluating the paper draft. We divide the scientific research lifecycle into three main components: Few-shot Scientific Knowledge Acquisition Scientific knowledge acquisition is the foundation of many downstream tasks. During the COVID-19 pandemic, we propose the COVID-KG to extract fine-grained multimedia knowledge graphs (KGs) from scientific papers for drug repurposing reports. However, fine-grained information extraction systems usually require large amounts of expert-annotated examples to perform effectively. Therefore, creating methods that can learn effectively from only a few examples, known as few-shot approaches, becomes vital in scientific knowledge acquisition. To address this limitation, we propose to utilize the knowledge consistency between input text and output knowledge elements for few-shot fine-grained scientific entity extraction. Integrating Domain Knowledge with Scientific LLM Reasoning Based on the scientific KGs in the previous component, we investigate scientific hypothesis generation, as well as experiment planning and execution. Simulating the human research process, we propose to augment LLMs with external heterogeneous KGs from previous papers as \"inspirations\" to generate novel scientific hypotheses and iteratively boosting the novelty of the generated hypothesis. Furthermore, by extending the current hypothesis generation approach into the biochemical domain, we retrieve relevant contextual information from multiple databases and propose a new enzyme sequence generation benchmark based on the given chemical reactions. Finally, to \\textit{incorporate external knowledge from both vision and text}, we introduce a new \\textit{multimedia procedure learning framework} to produce \\textit{visually trackable, inductive, and diverse} task scripts. Explainable Scientific Knowledge Dissemination To communicate a new idea to readers clearly and faithfully, evaluating the paper's quality is crucial to prevent distorted scientific dissemination. Therefore, we build an \\textit{explainable paper review generation system} to generate explainable review scores and comments, along with detailed evidence, based on KGs and papers. Beyond the paper text, tables concisely and clearly present complex information, facilitating comparisons and enhancing readability. To improve scientific table reasoning, we propose to decompose complex table reasoning tasks into a high-granularity atomic skill set. We further propose an incremental training procedure to ensure accurate information alignment in the reasoning procedure rather than indiscriminate connection to all available contexts. This work on AI4Scientist aims to open doors to the autonomous scientific research lifecycle, by equipping machines with both structured and unstructured knowledge from previous literature and enabling reasoning across different knowledge modalities. By automating the scientific research lifecycle, the \\textbf{AI4Scientist} can accelerate scientific discovery and reduce the time and resources required for research."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["AI4Scientist: Accelerating and democratizing scientific research lifecycle"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Han, Jiawei","Hakkani-Tur, Dilek","Zhao, Han","Neubig, Graham","Hope, Tom"],"dc:creator":["Wang, Qingyun"],"dc:date":["2025-04-15","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Qingyun Wang, accepted the attached license on 2025-04-11 at 19:28.","The student, Qingyun Wang, submitted this Dissertation for approval on 2025-04-11 at 19:28.","This Dissertation was approved for publication on 2025-04-15 at 09:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21750 on 2025-10-19 at 18:09:19","Millions of scientific papers are published annually, resulting in an information overload. Beyond this, scientific papers, known as \"Sleeping beauties\", sometimes remain largely unnoticed for long periods before suddenly attracting great attention. Moreover, the process of discovering new scientific hypotheses has remained slow, expensive, and highly specialist-dependent, due to the increasingly complex experiments. There exists a pressing need to help scientists digest and evaluate relevant papers, facilitating scientific discovery. However, computer-human collaboration in scientific hypothesis discovery is still exploratory and lacks a unified framework for analyzing the relevant tasks. This thesis tackles the problem of automating scientific literature understanding and scientific discovery by proposing an AI4Scientist. Unlike AI4Science, which creates AI-driven solutions for scientific challenges, we focus on AI4Scientists to empower scientists with AI tools to enhance their research lifecycle. The recent advancements in large language models (LLMs) raise the prospect that they may be able to solve those problems. Despite their impressive progress, LLMs often fail to effectively incorporate domain-specific knowledge and support their generation with enough evidence. Additionally, the expert-curated databases represent only a small fraction of knowledge in the entire domain, due to high annotation cost. To address this issue and lower the entry barrier for interdisciplinary collaboration, we develop AI tools to accelerate the entire research lifecycle for scientists, from knowledge acquisition, hypothesis generation, multimedia procedure planning for experiment design, experiment execution, conduction to writing, and evaluating the paper draft. We divide the scientific research lifecycle into three main components: Few-shot Scientific Knowledge Acquisition Scientific knowledge acquisition is the foundation of many downstream tasks. During the COVID-19 pandemic, we propose the COVID-KG to extract fine-grained multimedia knowledge graphs (KGs) from scientific papers for drug repurposing reports. However, fine-grained information extraction systems usually require large amounts of expert-annotated examples to perform effectively. Therefore, creating methods that can learn effectively from only a few examples, known as few-shot approaches, becomes vital in scientific knowledge acquisition. To address this limitation, we propose to utilize the knowledge consistency between input text and output knowledge elements for few-shot fine-grained scientific entity extraction. Integrating Domain Knowledge with Scientific LLM Reasoning Based on the scientific KGs in the previous component, we investigate scientific hypothesis generation, as well as experiment planning and execution. Simulating the human research process, we propose to augment LLMs with external heterogeneous KGs from previous papers as \"inspirations\" to generate novel scientific hypotheses and iteratively boosting the novelty of the generated hypothesis. Furthermore, by extending the current hypothesis generation approach into the biochemical domain, we retrieve relevant contextual information from multiple databases and propose a new enzyme sequence generation benchmark based on the given chemical reactions. Finally, to \\textit{incorporate external knowledge from both vision and text}, we introduce a new \\textit{multimedia procedure learning framework} to produce \\textit{visually trackable, inductive, and diverse} task scripts. Explainable Scientific Knowledge Dissemination To communicate a new idea to readers clearly and faithfully, evaluating the paper's quality is crucial to prevent distorted scientific dissemination. Therefore, we build an \\textit{explainable paper review generation system} to generate explainable review scores and comments, along with detailed evidence, based on KGs and papers. Beyond the paper text, tables concisely and clearly present complex information, facilitating comparisons and enhancing readability. To improve scientific table reasoning, we propose to decompose complex table reasoning tasks into a high-granularity atomic skill set. We further propose an incremental training procedure to ensure accurate information alignment in the reasoning procedure rather than indiscriminate connection to all available contexts. This work on AI4Scientist aims to open doors to the autonomous scientific research lifecycle, by equipping machines with both structured and unstructured knowledge from previous literature and enabling reasoning across different knowledge modalities. By automating the scientific research lifecycle, the \\textbf{AI4Scientist} can accelerate scientific discovery and reduce the time and resources required for research."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129195"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Qingyun Wang"],"dc:subject":["Natrual language generation","Scientific information extraction","Scientific knowledge reasoning","Scientific knowledge dissemination","AI4Scientist"],"dc:title":["AI4Scientist: Accelerating and democratizing scientific research lifecycle"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}