{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110851"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110851","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Joint document-level information extraction","abstract":"Constructing knowledge graphs from unstructured text is an important task that is relevant to many domains. Recently neural models have been used to great effect in order to solve many information extraction tasks. However, there are still many challenges that need to be solved before our models can achieve a level of natural language understanding that could be comparable to human. In order to accomplish that, we need to create models that can be optimized to jointly perform various IE tasks on large volumes of text, while properly utilizing all of the available information. In addition, as our models get more complex it is important to focus on producing explainable predictions, so that the reasoning behind a specific extracted fact can be understood by human users of the model. As a step towards solving these challenges, we introduce two new document-level IE models. The first model is trained to jointly perform identification, coreference, and classification of entities and events within a document by utilizing aggregated contextual information from each relevant mention. The second model builds on the first to extract relations with evidence between the entities in a document. We evaluate our models on the ACE-05+ and DocRed datasets respectively, and find improvements over the current SOTA in terms of F-score on entity, event, and evidence extraction.","abstract_html":"Constructing knowledge graphs from unstructured text is an important task that is relevant to many domains. Recently neural models have been used to great effect in order to solve many information extraction tasks. However, there are still many challenges that need to be solved before our models can achieve a level of natural language understanding that could be comparable to human. In order to accomplish that, we need to create models that can be optimized to jointly perform various IE tasks on large volumes of text, while properly utilizing all of the available information. In addition, as our models get more complex it is important to focus on producing explainable predictions, so that the reasoning behind a specific extracted fact can be understood by human users of the model. As a step towards solving these challenges, we introduce two new document-level IE models. The first model is trained to jointly perform identification, coreference, and classification of entities and events within a document by utilizing aggregated contextual information from each relevant mention. The second model builds on the first to extract relations with evidence between the entities in a document. We evaluate our models on the ACE-05+ and DocRed datasets respectively, and find improvements over the current SOTA in terms of F-score on entity, event, and evidence extraction.","abstract_has_math":false,"creators":["Kriman, Samuel"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T04:06:51Z","date_published":"2021-09-17T04:06:51Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Information Extraction","Entity","Relation","Event","Evidence","Joint Model"],"languages":["en"],"rights":["Copyright 2021 Samuel Kriman"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110851","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng"]},{"key":"dc:creator","label":"Author","values":["Kriman, Samuel"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T04:06:51Z","2023-09-17T04:07:01Z","2021-04-26","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Information Extraction","Entity","Relation","Event","Evidence","Joint Model"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Samuel Kriman"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110851"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Constructing knowledge graphs from unstructured text is an important task that is relevant to many domains. Recently neural models have been used to great effect in order to solve many information extraction tasks. However, there are still many challenges that need to be solved before our models can achieve a level of natural language understanding that could be comparable to human. In order to accomplish that, we need to create models that can be optimized to jointly perform various IE tasks on large volumes of text, while properly utilizing all of the available information. In addition, as our models get more complex it is important to focus on producing explainable predictions, so that the reasoning behind a specific extracted fact can be understood by human users of the model. As a step towards solving these challenges, we introduce two new document-level IE models. The first model is trained to jointly perform identification, coreference, and classification of entities and events within a document by utilizing aggregated contextual information from each relevant mention. The second model builds on the first to extract relations with evidence between the entities in a document. We evaluate our models on the ACE-05+ and DocRed datasets respectively, and find improvements over the current SOTA in terms of F-score on entity, event, and evidence extraction.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2023-05-01","The student, Samuel Kriman, accepted the attached license on 2021-04-23 at 09:57.","The student, Samuel Kriman, submitted this Thesis for approval on 2021-04-23 at 10:41.","This Thesis was approved for publication on 2021-04-26 at 15:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16525 on 2021-09-16 at 20:14:02","Made available in DSpace on 2021-09-17T04:06:51Z (GMT). No. of bitstreams: 2 KRIMAN-THESIS-2021.pdf: 2679757 bytes, checksum: cef2b40d7a940dcb4686159b8e9479b1 (MD5) LICENSE.txt: 4210 bytes, checksum: 95f42d594e9dd469ba6f196a24722239 (MD5) Previous issue date: 2021-04-26","Embargo set by: Seth Robbins for item 118697 Lift date: 2023-09-17T04:07:01Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Joint document-level information extraction"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng"],"dc:creator":["Kriman, Samuel"],"dc:date":["2021-09-17T04:06:51Z","2023-09-17T04:07:01Z","2021-04-26","2021-05"],"dc:description":["Constructing knowledge graphs from unstructured text is an important task that is relevant to many domains. Recently neural models have been used to great effect in order to solve many information extraction tasks. However, there are still many challenges that need to be solved before our models can achieve a level of natural language understanding that could be comparable to human. In order to accomplish that, we need to create models that can be optimized to jointly perform various IE tasks on large volumes of text, while properly utilizing all of the available information. In addition, as our models get more complex it is important to focus on producing explainable predictions, so that the reasoning behind a specific extracted fact can be understood by human users of the model. As a step towards solving these challenges, we introduce two new document-level IE models. The first model is trained to jointly perform identification, coreference, and classification of entities and events within a document by utilizing aggregated contextual information from each relevant mention. The second model builds on the first to extract relations with evidence between the entities in a document. We evaluate our models on the ACE-05+ and DocRed datasets respectively, and find improvements over the current SOTA in terms of F-score on entity, event, and evidence extraction.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2023-05-01","The student, Samuel Kriman, accepted the attached license on 2021-04-23 at 09:57.","The student, Samuel Kriman, submitted this Thesis for approval on 2021-04-23 at 10:41.","This Thesis was approved for publication on 2021-04-26 at 15:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16525 on 2021-09-16 at 20:14:02","Made available in DSpace on 2021-09-17T04:06:51Z (GMT). No. of bitstreams: 2 KRIMAN-THESIS-2021.pdf: 2679757 bytes, checksum: cef2b40d7a940dcb4686159b8e9479b1 (MD5) LICENSE.txt: 4210 bytes, checksum: 95f42d594e9dd469ba6f196a24722239 (MD5) Previous issue date: 2021-04-26","Embargo set by: Seth Robbins for item 118697 Lift date: 2023-09-17T04:07:01Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110851"],"dc:language":["en"],"dc:rights":["Copyright 2021 Samuel Kriman"],"dc:subject":["Information Extraction","Entity","Relation","Event","Evidence","Joint Model"],"dc:title":["Joint document-level information extraction"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}