University of Illinois at Urbana-Champaign
emrQA: A large corpus for question answering on electronic medical records
Abstract
dc:descriptionWe propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2019
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Pampari, Anusri
- Contributors dc:contributor
-
- Peng, Jian
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/102500
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/102500