Back to results

Iowa State University - Thesis & Dissertation

Consistency-aware and LLM-assisted methods for named entity recognition

Abstract

dc:description.abstract

Named Entity Recognition (NER) is a fundamental task in natural language processing and serves as a critical component for many downstream applications, including information extraction, biomedical text mining, and knowledge graph construction. Despite significant progress with neural models, most existing NER approaches process sentences independently, which can lead to inconsistent predictions for repeated entity mentions across a document. Moreover, the effectiveness of state-of-the-art NER systems, particularly in biomedical and clinical domains, is often constrained by the limited availability of high-quality labeled data, where annotation is costly and restricted by privacy concerns. The first part of this thesis focuses on improving document-level consistency in NER. While prior document-level NER methods attempt to share token-level representations across a document, such strategies may introduce noise when identical tokens appear in different semantic contexts. To address this limitation, we propose ScdNER, a span-based, consistency-aware document-level NER framework. ScdNER performs contextual feature fusion at the span level rather than the token level, enabling more precise sharing of information among repeated entity mentions. The model adopts a two-stage prediction framework: a binary classifier first identifies candidate entity spans, followed by a span-based key-value memory mechanism that selectively fuses global contextual features while suppressing noise from non-entity spans. Extensive experiments demonstrate that ScdNER achieves more consistent and accurate document-level predictions compared to existing approaches. Building on this foundation, the second part of the thesis investigates how large language models (LLMs) can be leveraged to enhance NER and relation extraction in data-scarce clinical settings. Although LLMs are capable of generating diverse and contextually rich text, their direct application to clinical information extraction poses challenges related to domain specificity, consistency, and long-document processing. To address these issues, we develop a structured LLM-based data augmentation framework that generates clinically faithful synthetic annotations, which are then used to enhance a modified BERT-based extraction model. A segmentation-based strategy combined with BiLSTM integration is introduced to preserve global context in long clinical notes. Experiments on public and proprietary datasets demonstrate substantial performance improvements, highlighting the effectiveness of structured LLM augmentation for clinical information extraction. The third part of this thesis further explores the role of LLMs in enhancing BERT-based biomedical NER models. While LLMs exhibit strong zero-shot and few-shot capabilities, they are often less efficient and less precise than fine-tuned BERT models for exact entity localization. Rather than replacing BERT, we investigate how LLMs can serve as auxiliary annotators to generate structured training data that complements human-labeled datasets. Through systematic evaluation, we show that LLM-generated annotations can effectively improve BERT performance, even when the auxiliary labels are noisy, demonstrating that BERT models are capable of denoising and generalizing from imperfect supervision. The final part of the thesis extends this direction by leveraging LLMs to generate fully synthetic training documents along with corresponding entity annotations. We explore fine-tuning LLMs using a combination of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to improve their ability to produce task-specific, high-quality annotated data. Inspired by recent advances in reinforcement learning for language models, we adopt a response-level advantage estimation strategy with a truncated objective to stabilize training while avoiding the need for an explicit value function. The resulting synthetic datasets are shown to be effective for training downstream NER models, further reducing reliance on expert annotations. Together, these studies present a unified framework for consistency-aware and LLM-enhanced named entity recognition, demonstrating how structured modeling and large language models can be combined to improve robustness, data efficiency, and scalability in biomedical and clinical NER systems.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
dissertation
Discipline thesis:degree_discipline
Computer science
Grantor
Iowa State University - Thesis & Dissertation
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wei, Ying
Advisor dc:contributor.advisor
  • Li, Qi

Subjects

dc:subject × 1

Rights

Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://dr.lib.iastate.edu/handle/20.500.12876/106602
OAI identifier oai:identifier
oai:dr.lib.iastate.edu:20.500.12876/106602

Chain of custody

source
Harvested from
Iowa State University
Base URL
dr.lib.iastate.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Wei, Ying. Consistency-aware and LLM-assisted methods for named entity recognition. dissertation thesis, Iowa State University - Thesis & Dissertation, 2026. https://dr.lib.iastate.edu/handle/20.500.12876/106602