{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127174"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127174","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Structure-enhanced text mining for science","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-03-28 without embargo terms","abstract_has_math":false,"creators":["Zhang, Yu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Tong, Hanghang","Wang, Wei","Shen, Zhihong","Abdelzaher, Tarek"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-11-04","date_published":"2024-11-04","updated_at":"2026-07-22T22:25:03Z","subjects":["Text Mining","Natural Language Processing","Artificial Intelligence For Science"],"languages":["en","eng"],"rights":["Copyright 2024 Yu Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127174","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Tong, Hanghang","Wang, Wei","Shen, Zhihong","Abdelzaher, Tarek"]},{"key":"dc:creator","label":"Author","values":["Zhang, Yu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-11-04","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Text Mining","Natural Language Processing","Artificial Intelligence For Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Yu Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127174"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Yu Zhang, accepted the attached license on 2024-10-30 at 21:44.","The student, Yu Zhang, submitted this Dissertation for approval on 2024-10-30 at 21:57.","This Dissertation was approved for publication on 2024-11-04 at 16:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21283 on 2025-03-28 at 14:25:22","Language models pre-trained on large scientific text corpora (e.g., research papers, electronic health records, and encyclopedia articles) have achieved remarkable success in many scientific text mining tasks. Meanwhile, text is usually accompanied by various types of structural signals, such as paper metadata, concept ontologies, and citation networks, that can potentially benefit scientific literature understanding. To enhance the effectiveness of scientific text mining methods, my doctoral research focuses on teaching language models to exploit structural information for fundamental and advanced domain-specific applications in science, with an emphasis on understanding and augmenting scientific discovery. This thesis summarizes my endeavors spanning the following three key subtopics: 1. Structure-Aware Fine-Grained Scientific Paper Classification. Automatically indexing scientific papers in a multi-dimensional and multi-granularity topic space not only facilitates flexible bibliographic exploration and analysis but also benefits a wide range of scientific applications. I have developed effective and efficient approaches that perform multi-label paper classification in an extremely large and multi-faceted label space (e.g., with 10,000-100,000 fields-of-study) by utilizing paper metadata, label taxonomy, heterogeneous information networks, and structures within full-text articles. 2. Structure-Aware Scientific Topic Discovery. Mining textual and structural topic-indicative signals (e.g., keywords, entities, metadata, and their combinations) has crucial applications in scientific discovery, such as finding biomarker proteins for different disease categories. I have proposed performant approaches for extracting label-indicative keywords and structural information, which enables the retrieval and synthesis of pseudo-labeled training samples and significantly enriches supervision in zero-shot and weakly supervised text mining. 3. Structure-Enhanced Language Model Pre-training for Scientific Applications. Scientific pre-trained language models can work alongside humans throughout the scientific discovery process by detecting and explaining relevant literature, unveiling knowledge structures, and evaluating research outcomes. I have systematically explored how to integrate fundamental scientific text mining tasks (e.g., paper classification, citation prediction, and literature retrieval) into language model pre-training to facilitate sophisticated applications, such as patient-to-article matching and peer review assignment. These efforts collectively pave the way for intelligent structure-enhanced text mining frameworks that process, utilize, and analyze scientific text data to accelerate science and innovation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Structure-enhanced text mining for science"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Tong, Hanghang","Wang, Wei","Shen, Zhihong","Abdelzaher, Tarek"],"dc:creator":["Zhang, Yu"],"dc:date":["2024-11-04","2024-12"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms","The student, Yu Zhang, accepted the attached license on 2024-10-30 at 21:44.","The student, Yu Zhang, submitted this Dissertation for approval on 2024-10-30 at 21:57.","This Dissertation was approved for publication on 2024-11-04 at 16:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21283 on 2025-03-28 at 14:25:22","Language models pre-trained on large scientific text corpora (e.g., research papers, electronic health records, and encyclopedia articles) have achieved remarkable success in many scientific text mining tasks. Meanwhile, text is usually accompanied by various types of structural signals, such as paper metadata, concept ontologies, and citation networks, that can potentially benefit scientific literature understanding. To enhance the effectiveness of scientific text mining methods, my doctoral research focuses on teaching language models to exploit structural information for fundamental and advanced domain-specific applications in science, with an emphasis on understanding and augmenting scientific discovery. This thesis summarizes my endeavors spanning the following three key subtopics: 1. Structure-Aware Fine-Grained Scientific Paper Classification. Automatically indexing scientific papers in a multi-dimensional and multi-granularity topic space not only facilitates flexible bibliographic exploration and analysis but also benefits a wide range of scientific applications. I have developed effective and efficient approaches that perform multi-label paper classification in an extremely large and multi-faceted label space (e.g., with 10,000-100,000 fields-of-study) by utilizing paper metadata, label taxonomy, heterogeneous information networks, and structures within full-text articles. 2. Structure-Aware Scientific Topic Discovery. Mining textual and structural topic-indicative signals (e.g., keywords, entities, metadata, and their combinations) has crucial applications in scientific discovery, such as finding biomarker proteins for different disease categories. I have proposed performant approaches for extracting label-indicative keywords and structural information, which enables the retrieval and synthesis of pseudo-labeled training samples and significantly enriches supervision in zero-shot and weakly supervised text mining. 3. Structure-Enhanced Language Model Pre-training for Scientific Applications. Scientific pre-trained language models can work alongside humans throughout the scientific discovery process by detecting and explaining relevant literature, unveiling knowledge structures, and evaluating research outcomes. I have systematically explored how to integrate fundamental scientific text mining tasks (e.g., paper classification, citation prediction, and literature retrieval) into language model pre-training to facilitate sophisticated applications, such as patient-to-article matching and peer review assignment. These efforts collectively pave the way for intelligent structure-enhanced text mining frameworks that process, utilize, and analyze scientific text data to accelerate science and innovation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127174"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Yu Zhang"],"dc:subject":["Text Mining","Natural Language Processing","Artificial Intelligence For Science"],"dc:title":["Structure-enhanced text mining for science"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:03Z"}