Back to results

University of Illinois at Urbana-Champaign

Granular text classification for biomedical natural language processing

Abstract

dc:description

Text classification is a classic NLP problem with numerous applications and use cases in sentiment analysis, spam detection, document organization, and information retrieval systems. While text classification techniques can be applied to documents of different lengths, supervised machine learning approaches for text classification require labeled documents similar in length to those that will be used at inference. This dissertation examines how labeled documents at higher levels of linguistic granularity (i.e., longer documents) may be synthesized to develop text classifiers at a lower level of linguistic granularity (i.e., shorter text). More specifically, we focus on Biomedical Natural Language Processing as our target domain and address (1) how well document-level classifiers perform for sentence-level classification; (2) how document-level labeled data may be synthesized to perform sentence-level text classification; and (3) how the performance of synthesized approaches compares against benchmark performances using labeled data. We present feature contribution analysis experiments in Naïve Bayes classifiers as well as self-attention experiments in BERT, and report sentence classification results that beat baseline performance for both model types. Finally, we present an extensive qualitative error analysis to identify major error trends in our results and discuss the significance of each error category with respect to our primary research questions.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Linguistics
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Abdar, Omid
Contributors dc:contributor
  • Schwartz, Lane
  • Stevens, Jon
  • Markee, Numa
  • Sadler, Randall
  • Ionin, Tania

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Omid Abdar
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/122213

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Abdar, Omid. Granular text classification for biomedical natural language processing. Dissertation thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/122213