Back to results

University of Illinois Urbana-Champaign

Towards knowledgeable natural language generation

Abstract

dc:description

Natural Language Generation (NLG) is a fundamental task in Natural Language Processing that has been studied for decades. NLG enables the automatic generation of coherent and contextually appropriate text. The recent advent of large language models (LLMs) has further advanced the field, but challenges such as factual correctness, coherence, and adaptability to domain-specific contexts persist. This thesis explores knowledge-enhanced NLG by integrating external knowledge and processing training data to improve generation quality and domain adaptability. The first research direction investigates the incorporation of structured external knowledge into NLG models. A key application is the generation of citing sentences in academic writing, where traditional models struggle to capture the complex motivations behind citations. To address this, we propose BACO, a framework that integrates background knowledge from citation networks and textual content from citing and cited papers. BACO enhances citation intent alignment and improves the coherence and informativeness of generated citing sentences. Another application of including external knowledge focuses on follow-up question generation in conversational surveys, where knowledge-driven approaches improve the relevance, diversity, and clarity of questions. A new dataset is introduced for this task, along with a two-stage generative model and a novel set of Gricean-inspired evaluation metrics. The second research direction explores leveraging training data as an additional knowledge source to enhance NLG performance. In abstractive summarization, we analyze the impact of extractivity on model performance and introduce a framework that mitigates excessive copying through an integer linear programming-based optimization technique. Additionally, we develop an approach that integrates a copy mechanism into BART with an auxiliary loss function, improving summary faithfulness and generalization. In extractive summarization, we propose a fine-grained autoregressive method that utilizes semantic tuples derived from predicate-argument structures to improve content selection. This approach outperforms traditional sentence-level classification methods and addresses redundancy and fixed-length cutoff issues. By advancing methodologies in knowledge-aware NLG, this thesis contributes to the broader goal of enhancing automatic text generation with informed, accurate, and domain-sensitive content. The findings provide insights into improving model training, refining generation quality, and expanding the applicability of NLG models across diverse tasks.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ge, Yubin
Contributors dc:contributor
  • Diesner, Jana
  • Ji, Heng
  • Hakkani-Tur, Dilek
  • Zhao, Han
  • Hazarika, Devamanyu

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Yubin Ge
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129469

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Ge, Yubin. Towards knowledgeable natural language generation. Dissertation thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129469