{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129469"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129469","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards knowledgeable natural language generation","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Ge, Yubin"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Diesner, Jana","Ji, Heng","Hakkani-Tur, Dilek","Zhao, Han","Hazarika, Devamanyu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-01","date_published":"2025-05-01","updated_at":"2026-07-22T22:25:05Z","subjects":["Natural Language Generation","Knowledgeable Text Generation"],"languages":["en","eng"],"rights":["Copyright 2025 Yubin Ge"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129469","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Diesner, Jana","Ji, Heng","Hakkani-Tur, Dilek","Zhao, Han","Hazarika, Devamanyu"]},{"key":"dc:creator","label":"Author","values":["Ge, Yubin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05-01","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Generation","Knowledgeable Text Generation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yubin Ge"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129469"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Yubin Ge, accepted the attached license on 2025-04-30 at 14:00.","The student, Yubin Ge, submitted this Dissertation for approval on 2025-04-30 at 14:11.","This Dissertation was approved for publication on 2025-05-01 at 06:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22122 on 2025-10-19 at 18:19:38","Natural Language Generation (NLG) is a fundamental task in Natural Language Processing that has been studied for decades. NLG enables the automatic generation of coherent and contextually appropriate text. The recent advent of large language models (LLMs) has further advanced the field, but challenges such as factual correctness, coherence, and adaptability to domain-specific contexts persist. This thesis explores knowledge-enhanced NLG by integrating external knowledge and processing training data to improve generation quality and domain adaptability. The first research direction investigates the incorporation of structured external knowledge into NLG models. A key application is the generation of citing sentences in academic writing, where traditional models struggle to capture the complex motivations behind citations. To address this, we propose BACO, a framework that integrates background knowledge from citation networks and textual content from citing and cited papers. BACO enhances citation intent alignment and improves the coherence and informativeness of generated citing sentences. Another application of including external knowledge focuses on follow-up question generation in conversational surveys, where knowledge-driven approaches improve the relevance, diversity, and clarity of questions. A new dataset is introduced for this task, along with a two-stage generative model and a novel set of Gricean-inspired evaluation metrics. The second research direction explores leveraging training data as an additional knowledge source to enhance NLG performance. In abstractive summarization, we analyze the impact of extractivity on model performance and introduce a framework that mitigates excessive copying through an integer linear programming-based optimization technique. Additionally, we develop an approach that integrates a copy mechanism into BART with an auxiliary loss function, improving summary faithfulness and generalization. In extractive summarization, we propose a fine-grained autoregressive method that utilizes semantic tuples derived from predicate-argument structures to improve content selection. This approach outperforms traditional sentence-level classification methods and addresses redundancy and fixed-length cutoff issues. By advancing methodologies in knowledge-aware NLG, this thesis contributes to the broader goal of enhancing automatic text generation with informed, accurate, and domain-sensitive content. The findings provide insights into improving model training, refining generation quality, and expanding the applicability of NLG models across diverse tasks."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards knowledgeable natural language generation"]}]}],"canonical_facts":{"dc:contributor":["Diesner, Jana","Ji, Heng","Hakkani-Tur, Dilek","Zhao, Han","Hazarika, Devamanyu"],"dc:creator":["Ge, Yubin"],"dc:date":["2025-05-01","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Yubin Ge, accepted the attached license on 2025-04-30 at 14:00.","The student, Yubin Ge, submitted this Dissertation for approval on 2025-04-30 at 14:11.","This Dissertation was approved for publication on 2025-05-01 at 06:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22122 on 2025-10-19 at 18:19:38","Natural Language Generation (NLG) is a fundamental task in Natural Language Processing that has been studied for decades. NLG enables the automatic generation of coherent and contextually appropriate text. The recent advent of large language models (LLMs) has further advanced the field, but challenges such as factual correctness, coherence, and adaptability to domain-specific contexts persist. This thesis explores knowledge-enhanced NLG by integrating external knowledge and processing training data to improve generation quality and domain adaptability. The first research direction investigates the incorporation of structured external knowledge into NLG models. A key application is the generation of citing sentences in academic writing, where traditional models struggle to capture the complex motivations behind citations. To address this, we propose BACO, a framework that integrates background knowledge from citation networks and textual content from citing and cited papers. BACO enhances citation intent alignment and improves the coherence and informativeness of generated citing sentences. Another application of including external knowledge focuses on follow-up question generation in conversational surveys, where knowledge-driven approaches improve the relevance, diversity, and clarity of questions. A new dataset is introduced for this task, along with a two-stage generative model and a novel set of Gricean-inspired evaluation metrics. The second research direction explores leveraging training data as an additional knowledge source to enhance NLG performance. In abstractive summarization, we analyze the impact of extractivity on model performance and introduce a framework that mitigates excessive copying through an integer linear programming-based optimization technique. Additionally, we develop an approach that integrates a copy mechanism into BART with an auxiliary loss function, improving summary faithfulness and generalization. In extractive summarization, we propose a fine-grained autoregressive method that utilizes semantic tuples derived from predicate-argument structures to improve content selection. This approach outperforms traditional sentence-level classification methods and addresses redundancy and fixed-length cutoff issues. By advancing methodologies in knowledge-aware NLG, this thesis contributes to the broader goal of enhancing automatic text generation with informed, accurate, and domain-sensitive content. The findings provide insights into improving model training, refining generation quality, and expanding the applicability of NLG models across diverse tasks."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129469"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Yubin Ge"],"dc:subject":["Natural Language Generation","Knowledgeable Text Generation"],"dc:title":["Towards knowledgeable natural language generation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}