{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/102500"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/102500","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"emrQA: A large corpus for question answering on electronic medical records","abstract":"We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.","abstract_html":"We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.","abstract_has_math":false,"creators":["Pampari, Anusri"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-02-06T19:36:41Z","date_published":"2019-02-06T19:36:41Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Electronic Medical Records, Question Answering, Logical Forms, Semantic Parsing, Dataset Generation, Closed Domain, i2b2"],"languages":["en"],"rights":["Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/102500","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Pampari, Anusri"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-02-06T19:36:41Z","2018-12-11","2018-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electronic Medical Records, Question Answering, Logical Forms, Semantic Parsing, Dataset Generation, Closed Domain, i2b2"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/102500"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-02-05 without embargo terms","The student, Anusri Pampari, accepted the attached license on 2018-12-10 at 17:04.","The student, Anusri Pampari, submitted this Thesis for approval on 2018-12-10 at 17:14.","This Thesis was approved for publication on 2018-12-11 at 16:32.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13258 on 2019-02-05 at 11:15:51","Made available in DSpace on 2019-02-06T19:36:41Z (GMT). No. of bitstreams: 2 PAMPARI-THESIS-2018.pdf: 647132 bytes, checksum: 7f5c13ac50da2e60244c6e0cc573fc40 (MD5) LICENSE.txt: 4211 bytes, checksum: 766b4aeffc00e3dd1bd91d2b58ec3638 (MD5) Previous issue date: 2018-12-11"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["emrQA: A large corpus for question answering on electronic medical records"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian"],"dc:creator":["Pampari, Anusri"],"dc:date":["2019-02-06T19:36:41Z","2018-12-11","2018-12"],"dc:description":["We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-02-05 without embargo terms","The student, Anusri Pampari, accepted the attached license on 2018-12-10 at 17:04.","The student, Anusri Pampari, submitted this Thesis for approval on 2018-12-10 at 17:14.","This Thesis was approved for publication on 2018-12-11 at 16:32.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13258 on 2019-02-05 at 11:15:51","Made available in DSpace on 2019-02-06T19:36:41Z (GMT). No. of bitstreams: 2 PAMPARI-THESIS-2018.pdf: 647132 bytes, checksum: 7f5c13ac50da2e60244c6e0cc573fc40 (MD5) LICENSE.txt: 4211 bytes, checksum: 766b4aeffc00e3dd1bd91d2b58ec3638 (MD5) Previous issue date: 2018-12-11"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/102500"],"dc:language":["en"],"dc:rights":["Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018"],"dc:subject":["Electronic Medical Records, Question Answering, Logical Forms, Semantic Parsing, Dataset Generation, Closed Domain, i2b2"],"dc:title":["emrQA: A large corpus for question answering on electronic medical records"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}