{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/46618"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/46618","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Information extraction for clinical narratives","abstract":"Recent US government initiatives have made available a large number of Electronic Health Records (EHRs). These EHRs contain valuable information which can be used in Clinical Decision Support (CDS). So, Information Extraction (IE) from EHRs is a very promising research area. In this thesis, I focus on two tasks namely Mention Detection and Coreference Resolution. A lot of domain knowledge is available regarding clinical narratives. There are also several tools like SpecialistLexicalTools, MetaMap, etc. which help in analyzing clinical narratives. I integrate the domain knowledge and features derived from these tools in the local statistical models. Clinical narratives have a very special format which gives several interconnections between the tasks of mention detection and coreference resolution. A joint formulation for these two tasks has been presented in this thesis. Along with this, there is also a discussion regarding joint formulation for finding the mention types together. Soft constraints have been used while formulating the inference tasks. Softening the constraints is helpful because it allows the constraints to be violated during inference. Joint formulation is based on the fact that only local models are learned in the training phase. Inconsistencies in the decisions based on local models are resolved during the global inference step. I report the best results, to date, on end-to-end coreference resolution. The joint formulation presented in this thesis is very general and would benefit other information extraction tasks as well. I have made the systems described in this thesis publicly available for research use.","abstract_html":"Recent US government initiatives have made available a large number of Electronic Health Records (EHRs). These EHRs contain valuable information which can be used in Clinical Decision Support (CDS). So, Information Extraction (IE) from EHRs is a very promising research area. In this thesis, I focus on two tasks namely Mention Detection and Coreference Resolution. A lot of domain knowledge is available regarding clinical narratives. There are also several tools like SpecialistLexicalTools, MetaMap, etc. which help in analyzing clinical narratives. I integrate the domain knowledge and features derived from these tools in the local statistical models. Clinical narratives have a very special format which gives several interconnections between the tasks of mention detection and coreference resolution. A joint formulation for these two tasks has been presented in this thesis. Along with this, there is also a discussion regarding joint formulation for finding the mention types together. Soft constraints have been used while formulating the inference tasks. Softening the constraints is helpful because it allows the constraints to be violated during inference. Joint formulation is based on the fact that only local models are learned in the training phase. Inconsistencies in the decisions based on local models are resolved during the global inference step. I report the best results, to date, on end-to-end coreference resolution. The joint formulation presented in this thesis is very general and would benefit other information extraction tasks as well. I have made the systems described in this thesis publicly available for research use.","abstract_has_math":false,"creators":["Jindal, Prateek"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Roth, Dan","Gunter, Carl A.","Zhai, ChengXiang","Chapman, Wendy"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-01-16T17:56:22Z","date_published":"2014-01-16T17:56:22Z","updated_at":"2026-07-22T22:25:36Z","subjects":["Natural Language Processing","Electronic Health Records","Mention Detection","Coreference Resolution","Drug Abuse Events","Set Expansion","Integer Linear Programming","Integer Quadratic Programming","Information Extraction","Text Mining","Clinical Narratives","Temporal Expression Extraction","Joint Inference"],"languages":["en"],"rights":["Copyright 2013 Prateek Jindal"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/46618","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roth, Dan","Gunter, Carl A.","Zhai, ChengXiang","Chapman, Wendy"]},{"key":"dc:creator","label":"Author","values":["Jindal, Prateek"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-01-16T17:56:22Z","2013-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing","Electronic Health Records","Mention Detection","Coreference Resolution","Drug Abuse Events","Set Expansion","Integer Linear Programming","Integer Quadratic Programming","Information Extraction","Text Mining","Clinical Narratives","Temporal Expression Extraction","Joint Inference"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Prateek Jindal"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/46618"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Recent US government initiatives have made available a large number of Electronic Health Records (EHRs). These EHRs contain valuable information which can be used in Clinical Decision Support (CDS). So, Information Extraction (IE) from EHRs is a very promising research area. In this thesis, I focus on two tasks namely Mention Detection and Coreference Resolution. A lot of domain knowledge is available regarding clinical narratives. There are also several tools like SpecialistLexicalTools, MetaMap, etc. which help in analyzing clinical narratives. I integrate the domain knowledge and features derived from these tools in the local statistical models. Clinical narratives have a very special format which gives several interconnections between the tasks of mention detection and coreference resolution. A joint formulation for these two tasks has been presented in this thesis. Along with this, there is also a discussion regarding joint formulation for finding the mention types together. Soft constraints have been used while formulating the inference tasks. Softening the constraints is helpful because it allows the constraints to be violated during inference. Joint formulation is based on the fact that only local models are learned in the training phase. Inconsistencies in the decisions based on local models are resolved during the global inference step. I report the best results, to date, on end-to-end coreference resolution. The joint formulation presented in this thesis is very general and would benefit other information extraction tasks as well. I have made the systems described in this thesis publicly available for research use.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-09-26T18:09:13Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Jindal_Prateek_latex.zip: 3311882 bytes, checksum: ede17a21e0f98cf9dca28d676c661b88 (MD5) Jindal_Prateek.pdf: 665238 bytes, checksum: 536d6f0b09c7707871db6013f42f70b2 (MD5)","Made available in DSpace on 2014-01-16T17:56:22Z (GMT). No. of bitstreams: 3 Prateek_Jindal.pdf: 665238 bytes, checksum: 536d6f0b09c7707871db6013f42f70b2 (MD5) Jindal_Prateek_latex.zip: 3311882 bytes, checksum: ede17a21e0f98cf9dca28d676c661b88 (MD5) license.txt: 4063 bytes, checksum: feaf65aa170cdb985b39882dab27832f (MD5)"]},{"key":"dc:title","label":"Title","values":["Information extraction for clinical narratives"]}]}],"canonical_facts":{"dc:contributor":["Roth, Dan","Gunter, Carl A.","Zhai, ChengXiang","Chapman, Wendy"],"dc:creator":["Jindal, Prateek"],"dc:date":["2014-01-16T17:56:22Z","2013-12"],"dc:description":["Recent US government initiatives have made available a large number of Electronic Health Records (EHRs). These EHRs contain valuable information which can be used in Clinical Decision Support (CDS). So, Information Extraction (IE) from EHRs is a very promising research area. In this thesis, I focus on two tasks namely Mention Detection and Coreference Resolution. A lot of domain knowledge is available regarding clinical narratives. There are also several tools like SpecialistLexicalTools, MetaMap, etc. which help in analyzing clinical narratives. I integrate the domain knowledge and features derived from these tools in the local statistical models. Clinical narratives have a very special format which gives several interconnections between the tasks of mention detection and coreference resolution. A joint formulation for these two tasks has been presented in this thesis. Along with this, there is also a discussion regarding joint formulation for finding the mention types together. Soft constraints have been used while formulating the inference tasks. Softening the constraints is helpful because it allows the constraints to be violated during inference. Joint formulation is based on the fact that only local models are learned in the training phase. Inconsistencies in the decisions based on local models are resolved during the global inference step. I report the best results, to date, on end-to-end coreference resolution. The joint formulation presented in this thesis is very general and would benefit other information extraction tasks as well. I have made the systems described in this thesis publicly available for research use.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-09-26T18:09:13Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Jindal_Prateek_latex.zip: 3311882 bytes, checksum: ede17a21e0f98cf9dca28d676c661b88 (MD5) Jindal_Prateek.pdf: 665238 bytes, checksum: 536d6f0b09c7707871db6013f42f70b2 (MD5)","Made available in DSpace on 2014-01-16T17:56:22Z (GMT). No. of bitstreams: 3 Prateek_Jindal.pdf: 665238 bytes, checksum: 536d6f0b09c7707871db6013f42f70b2 (MD5) Jindal_Prateek_latex.zip: 3311882 bytes, checksum: ede17a21e0f98cf9dca28d676c661b88 (MD5) license.txt: 4063 bytes, checksum: feaf65aa170cdb985b39882dab27832f (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/46618"],"dc:language":["en"],"dc:rights":["Copyright 2013 Prateek Jindal"],"dc:subject":["Natural Language Processing","Electronic Health Records","Mention Detection","Coreference Resolution","Drug Abuse Events","Set Expansion","Integer Linear Programming","Integer Quadratic Programming","Information Extraction","Text Mining","Clinical Narratives","Temporal Expression Extraction","Joint Inference"],"dc:title":["Information extraction for clinical narratives"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:36Z"}