Iowa State University - Thesis & Dissertation
Consistency-aware and LLM-assisted methods for named entity recognition
Abstract
dc:description.abstractNamed Entity Recognition (NER) is a fundamental task in natural language processing and serves as a critical component for many downstream applications, including information extraction, biomedical text mining, and knowledge graph construction. Despite significant progress with neural models, most existing NER approaches process sentences independently, which can lead to inconsistent predictions for repeated entity mentions across a document. Moreover, the effectiveness of state-of-the-art NER systems, particularly in biomedical and clinical domains, is often constrained by the limited availability of high-quality labeled data, where annotation is costly and restricted by privacy concerns. The first part of this thesis focuses on improving document-level consistency in NER. While prior document-level NER methods attempt to share token-level representations across a document, such strategies may introduce noise when identical tokens appear in different semantic contexts. To address this limitation, we propose ScdNER, a span-based, consistency-aware document-level NER framework. ScdNER performs contextual feature fusion at the span level rather than the token level, enabling more precise sharing of information among repeated entity mentions. The model adopts a two-stage prediction framework: a binary classifier first identifies candidate entity spans, followed by a span-based key-value memory mechanism that selectively fuses global contextual features while suppressing noise from non-entity spans. Extensive experiments demonstrate that ScdNER achieves more consistent and accurate document-level predictions compared to existing approaches. Building on this foundation, the second part of the thesis investigates how large language models (LLMs) can be leveraged to enhance NER and relation extraction in data-scarce clinical settings. Although LLMs are capable of generating diverse and contextually rich text, their direct application to clinical information extraction poses challenges related to domain specificity, consistency, and long-document processing. To address these issues, we develop a structured LLM-based data augmentation framework that generates clinically faithful synthetic annotations, which are then used to enhance a modified BERT-based extraction model. A segmentation-based strategy combined with BiLSTM integration is introduced to preserve global context in long clinical notes. Experiments on public and proprietary datasets demonstrate substantial performance improvements, highlighting the effectiveness of structured LLM augmentation for clinical information extraction. The third part of this thesis further explores the role of LLMs in enhancing BERT-based biomedical NER models. While LLMs exhibit strong zero-shot and few-shot capabilities, they are often less efficient and less precise than fine-tuned BERT models for exact entity localization. Rather than replacing BERT, we investigate how LLMs can serve as auxiliary annotators to generate structured training data that complements human-labeled datasets. Through systematic evaluation, we show that LLM-generated annotations can effectively improve BERT performance, even when the auxiliary labels are noisy, demonstrating that BERT models are capable of denoising and generalizing from imperfect supervision. The final part of the thesis extends this direction by leveraging LLMs to generate fully synthetic training documents along with corresponding entity annotations. We explore fine-tuning LLMs using a combination of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to improve their ability to produce task-specific, high-quality annotated data. Inspired by recent advances in reinforcement learning for language models, we adopt a response-level advantage estimation strategy with a truncated objective to stabilize training while avoiding the need for an explicit value function. The resulting synthetic datasets are shown to be effective for training downstream NER models, further reducing reliance on expert annotations. Together, these studies present a unified framework for consistency-aware and LLM-enhanced named entity recognition, demonstrating how structured modeling and large language models can be combined to improve robustness, data efficiency, and scalability in biomedical and clinical NER systems.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Level thesis:degree_level
- dissertation
- Discipline thesis:degree_discipline
- Computer science
- Grantor
- Iowa State University - Thesis & Dissertation
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Wei, Ying
- Advisor dc:contributor.advisor
-
- Li, Qi
Subjects
dc:subject × 1Rights
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://dr.lib.iastate.edu/handle/20.500.12876/106602
- OAI identifier oai:identifier
- oai:dr.lib.iastate.edu:20.500.12876/106602