Back to results

University of Illinois at Urbana-Champaign

Towards fine-grained automated age extraction for precision medicine

Abstract

dc:description

Precision medicine (PM) aims to customize interventions for individuals. Although genetic factors have received significant attention, non-genetic population characteristics such as age, race, and gender remain underutilized due to challenges in extracting information at the right level of specificity from the literature. Our goal is to automatically identify age, where authors use either narrative (e.g., adult, child) or numbers where the latter is further delineated into mean, median, minimum, maximum (e.g., 20 and 60 respectively from the text aged 20-60), standard deviation, and upper and lower bounds (e.g., >60, less than 50). To address this issue, we conducted 2 experiments. In the first experiment, we used three information extraction methods (rule-based, extractive question-answering, and conditional generation) to identify age at the level of detail needed by PM. We then conduct experiments using an existing evidence-based medicine natural language processing (EBM-NLP) dataset (that we have augmented to include the extra detail) and introduce a new dataset focused on breast cancer, where the age at diagnosis significantly impacts intervention possibilities. Extractive question answering consistently outperformed the other techniques, achieving the highest F1-scores of 0.988, 0.961, 0.991, 0.994, and 0.667 for mean, median, minimum, maximum, and standard deviation, respectively, in the breast cancer dataset. Our analysis also revealed 200 missing values in EBM-NLP and that training models for each facet rely heavily on having a substantial amount of training data and computational resources. The second experiment addresses this limitation, where we explore information extraction active learning. We curated a new breast cancer data set and conducted a comparative study on conditional random fields and bidirectional long short-term memory-conditional random field (Bi-LSTM-CRF) models and sampling techniques. Our findings reveal that the Bi-LSTM-CRF model, coupled with random sampling, outperforms other approaches, achieving F1-scores of 0.653, 0.719, 0.924, and 0.977 for mean, median, minimum, and maximum, respectively. The active learning approach did not work for standard deviation. Overall, these results demonstrate that automated methods can identify age and the latter approach is promising to identify other regarding non-genetic factors, for precision medicine. This thesis contributes to the fields of precision medicine and information science by providing publicly available datasets, including a new breast cancer dataset and a revised EBM-NLP dataset that includes the level of specificity required by PM, to enable others to hone their automated methods.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Information Management
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Salvi, Rohan Charudatt
Contributors dc:contributor
  • Blake, Catherine
  • Bosch, Nigel

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Rohan Charudatt Salvi
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/121277

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Salvi, Rohan Charudatt. Towards fine-grained automated age extraction for precision medicine. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/121277