Back to results

University of Illinois at Urbana-Champaign

emrQA: A large corpus for question answering on electronic medical records

Abstract

dc:description

We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2019

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pampari, Anusri
Contributors dc:contributor
  • Peng, Jian

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/102500
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/102500

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Pampari, Anusri. emrQA: A large corpus for question answering on electronic medical records. Thesis thesis, University of Illinois at Urbana-Champaign, 2019. http://hdl.handle.net/2142/102500