{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129870"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129870","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Natural language processing for supporting impact assessment of funded projects","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Han, Kanyao"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Information Sciences","degree_department":null,"school":null,"contributors":["Diesner, Jana","Schneider, Jodi","Kilicoglu, Halil","Miller, Daniel C."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-15","date_published":"2025-07-15","updated_at":"2026-07-22T22:25:06Z","subjects":["Natural Language Processing (nlp)","Impact Assessment","Text Mining","Information Extraction","Funder Name Disambiguation","Funding Analysis","Domain-specific Nlp","Low-resource Nlp"],"languages":["en","eng"],"rights":["Copyright 2025 Kanyao Han"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129870","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Diesner, Jana","Schneider, Jodi","Kilicoglu, Halil","Miller, Daniel C."]},{"key":"dc:creator","label":"Author","values":["Han, Kanyao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-15","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Information Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing (nlp)","Impact Assessment","Text Mining","Information Extraction","Funder Name Disambiguation","Funding Analysis","Domain-specific Nlp","Low-resource Nlp"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Kanyao Han"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129870"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Kanyao Han, accepted the attached license on 2025-07-14 at 23:45.","The student, Kanyao Han, submitted this Dissertation for approval on 2025-07-14 at 23:53.","This Dissertation was approved for publication on 2025-07-15 at 16:19.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22526 on 2025-10-20 at 16:57:52","Funding from organizations like the U.S. National Science Foundation plays a crucial role in supporting researchers and practitioners in advancing scientific knowledge, promoting societal progress, and protecting the environment, among other goals. As a result, both organizations and researchers are keen to understand how such funding is distributed across various projects and disciplines, as well as the outcomes and impacts generated by these projects. A comprehensive analysis of diverse text-based data sources that document funding allocations, research outcomes, and broader impacts can help deepen this understanding. These data sources include project reports submitted to funders as well as outcomes published in research articles. However, annotating and analyzing text-based data, even at moderate volumes, can be time-consuming and costly. Researchers must process lengthy and large-scale datasets to identify meaningful information for analysis. This dissertation aims to leverage computational methods, particularly from the fields of Natural Language Processing (NLP) and Machine Learning (ML), to assist researchers and practitioners in managing text-based data more efficiently and effectively. By automating or semi-automating processes such as information extraction, data cleaning, and classification, this work seeks to reduce the workload associated with data processing and annotation. This dissertation explores how NLP and ML techniques can be developed and used to handle data from social and scientific research under three challenging conditions: (1) disorganized, complex, lengthy, or incomplete datasets; (2) limited availability of annotated data; and (3) the need for domain-specific analysis schemas. By addressing these challenges, this dissertation aims to develop innovative approaches to aid in the analysis of funding allocation and the assessment of the impact of funded projects, with three studies being presented. First, analyzing past funding allocations can offer valuable insights into funding patterns in previous research. However, such analyses are often hindered by inconsistent and ambiguous naming conventions for funding organizations in publication records. This dissertation proposes a framework for fine-tuning a model to disambiguate funder names. Second, categorizing project reports can provide valuable insights into how funding is allocated across different project themes. Despite the availability of various categorization methods that typically require large volumes of annotated data for model fine-tuning or training, little is known about how to build effective models when: (a) categorizing texts demands substantial domain expertise and/or detailed reading; (b) only a limited number of annotated documents are available for training; and (c) no relevant computational resources, such as effective pre-trained models, exist. This dissertation introduces and evaluates a categorization method that combines expert knowledge with computational models to develop domain-specific categorization models. Third, with funding agencies increasingly demanding evidence of the social impact of scientific research, impact assessment has become critical. However, challenges remain in categorizing research reports due to the absence of a comprehensive impact classification schema and standardized reporting formats across domains. This dissertation addresses these gaps by developing and evaluating a classification schema for assessing the impact of funded research projects across domains, assisted by NLP and ML techniques. This dissertation advances knowledge by (1) developing novel frameworks for cleaning, annotating, and extracting valuable information from publication records and project reports; (2) providing insights into funding allocation in scientific research and biodiversity conservation; and (3) enhancing the understanding of the impacts described by funded projects."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Natural language processing for supporting impact assessment of funded projects"]}]}],"canonical_facts":{"dc:contributor":["Diesner, Jana","Schneider, Jodi","Kilicoglu, Halil","Miller, Daniel C."],"dc:creator":["Han, Kanyao"],"dc:date":["2025-07-15","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Kanyao Han, accepted the attached license on 2025-07-14 at 23:45.","The student, Kanyao Han, submitted this Dissertation for approval on 2025-07-14 at 23:53.","This Dissertation was approved for publication on 2025-07-15 at 16:19.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22526 on 2025-10-20 at 16:57:52","Funding from organizations like the U.S. National Science Foundation plays a crucial role in supporting researchers and practitioners in advancing scientific knowledge, promoting societal progress, and protecting the environment, among other goals. As a result, both organizations and researchers are keen to understand how such funding is distributed across various projects and disciplines, as well as the outcomes and impacts generated by these projects. A comprehensive analysis of diverse text-based data sources that document funding allocations, research outcomes, and broader impacts can help deepen this understanding. These data sources include project reports submitted to funders as well as outcomes published in research articles. However, annotating and analyzing text-based data, even at moderate volumes, can be time-consuming and costly. Researchers must process lengthy and large-scale datasets to identify meaningful information for analysis. This dissertation aims to leverage computational methods, particularly from the fields of Natural Language Processing (NLP) and Machine Learning (ML), to assist researchers and practitioners in managing text-based data more efficiently and effectively. By automating or semi-automating processes such as information extraction, data cleaning, and classification, this work seeks to reduce the workload associated with data processing and annotation. This dissertation explores how NLP and ML techniques can be developed and used to handle data from social and scientific research under three challenging conditions: (1) disorganized, complex, lengthy, or incomplete datasets; (2) limited availability of annotated data; and (3) the need for domain-specific analysis schemas. By addressing these challenges, this dissertation aims to develop innovative approaches to aid in the analysis of funding allocation and the assessment of the impact of funded projects, with three studies being presented. First, analyzing past funding allocations can offer valuable insights into funding patterns in previous research. However, such analyses are often hindered by inconsistent and ambiguous naming conventions for funding organizations in publication records. This dissertation proposes a framework for fine-tuning a model to disambiguate funder names. Second, categorizing project reports can provide valuable insights into how funding is allocated across different project themes. Despite the availability of various categorization methods that typically require large volumes of annotated data for model fine-tuning or training, little is known about how to build effective models when: (a) categorizing texts demands substantial domain expertise and/or detailed reading; (b) only a limited number of annotated documents are available for training; and (c) no relevant computational resources, such as effective pre-trained models, exist. This dissertation introduces and evaluates a categorization method that combines expert knowledge with computational models to develop domain-specific categorization models. Third, with funding agencies increasingly demanding evidence of the social impact of scientific research, impact assessment has become critical. However, challenges remain in categorizing research reports due to the absence of a comprehensive impact classification schema and standardized reporting formats across domains. This dissertation addresses these gaps by developing and evaluating a classification schema for assessing the impact of funded research projects across domains, assisted by NLP and ML techniques. This dissertation advances knowledge by (1) developing novel frameworks for cleaning, annotating, and extracting valuable information from publication records and project reports; (2) providing insights into funding allocation in scientific research and biodiversity conservation; and (3) enhancing the understanding of the impacts described by funded projects."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129870"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Kanyao Han"],"dc:subject":["Natural Language Processing (nlp)","Impact Assessment","Text Mining","Information Extraction","Funder Name Disambiguation","Funding Analysis","Domain-specific Nlp","Low-resource Nlp"],"dc:title":["Natural language processing for supporting impact assessment of funded projects"],"dc:type":["text"],"thesis:degree_discipline":["Information Sciences"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}