Back to results

University of Illinois at Urbana-Champaign

Automatic rare disease extraction based on large language models

Abstract

dc:description

Identifying and extracting information on rare diseases is crucial in various medical contexts. However, mature rare disease extraction methods are lacking in low-resource settings. In this paper, we aim to create an end-to-end system called AutoRD, which automates extracting information from clinical text about rare diseases. We achieve this using large language models and medical knowledge graphs developed from open-source medical ontologies. Large language models (LLMs) aid in language analysis, while knowledge graphs provide content-specific facts, thus filling in any information gaps. Our system, AutoRD, is a software pipeline involving data preprocessing, entity extraction, relation extraction, entity calibration, and knowledge graph construction. Large language models and open-source data are leveraged throughout the pipeline. We have conducted various tests to evaluate the performance of AutoRD and highlighted its strengths and limitations in this paper. We quantitatively evaluate our system in terms of entity extraction, relation extraction, and the performance of knowledge graph construction. AutoRD achieves an overall F1 score of 47.3%, an improvement of 0.8% compared to the fine-tuned model, and a 14.4% improvement compared to the base LLM. In detail, AutoRD achieves an overall entity extraction F1 score of 56.1% (rare_disease: 83.5%, disease: 35.8%, symptom_and_sign: 46.1%, anaphor: 67.5%) and an overall relation extraction F1 score of 38.6% (produces: 34.7%, increases_risk_of: 12.4%, is_a: 37.4%, is_acronym: 44.1%, is_synonym: 16.3%, anaphora: 57.5%). Our qualitative experiment also demonstrates that the performance in constructing the knowledge graph is commendable. Several designs, including the incorporation of ontologies-enhanced LLMs, contribute to the improvement of AutoRD. AutoRD demonstrates superior performance compared to other methods, demonstrating the potential of LLM applications in rare disease detection and AI for healthcare.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Cao, Lang
Contributors dc:contributor
  • Sun, Jimeng

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Lang Cao
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/124135

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Cao, Lang. Automatic rare disease extraction based on large language models. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/124135