University of Illinois - Chicago
Advancing NLP Frontiers in Information Extraction, Opinion Mining, and Text Synthesis
Abstract
dc:descriptionAdvances in generative AI and transformer architectures have propelled NLP toward systems capable of extracting, interpreting, and synthesizing information with greater accuracy and contextual awareness. This dissertation presents an integrated research program spanning three pillars: information extraction, opinion mining, and text synthesis, where each stage builds upon the previous to form a cohesive pipeline from distilling knowledge to generating coherent, faithful narratives. In the first pillar, we address keyphrase generation as an information extraction task for capturing the main topics and salient concepts of research papers. To overcome limited context and data scarcity, we introduce FullTextKP, the first large-scale full-text keyphrase dataset, and propose domain-agnostic augmentation methods for low-resource settings. These advances form the foundation for the second pillar, which targets opinion mining through the introduction of the Target-Stance Extraction (TSE) task to jointly identify targets and detect corresponding stances, even in noisy, real-world scenarios. Building on this, we develop Stanceformer, a target-aware transformer architecture that biases attention toward target terms, achieving state-of-the-art performance and robust cross-domain generalization. The third pillar extends these foundations to text synthesis via the Scientific Introduction Generation (SciIG) task and benchmark, which evaluate large language models on coherence, faithfulness, content coverage, and citation correctness in scholarly writing. A key innovation is the use of keyphrase-based coverage measures, originating from our first pillar, to assess whether generated introductions preserve essential information without distortion. Collectively, these contributions deliver datasets, architectures, and evaluation frameworks that advance context-aware, reliable, and scalable NLP systems, demonstrating the transformative potential of generative AI in scholarly and social domains.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Krishna Kant Garg (23291344)
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- In Copyright
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.31451095.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/31451095