Back to results

Texas A&M University

Connecting the Dots: Improving Information Extraction by Modeling the Non-Sequential Dependencies of Entity Mentions

Abstract

dc:description.abstract

Information Extraction (IE) aims at automatically extracting structured information from unstructured and semi-structured documents. It is an important and challenging topic in Natural Language Processing (NLP) that plays a critical role in downstream applications such as Question Answering and Summarizing. It includes many sub-tasks like Named Entity Recognition (NER), Relation Extraction, Event Extraction, Table Annotations, etc. Previous works on IE mainly focus on extracting knowledge from a document sentence by sentence and the text encoding models (e.g., RNNs, CNNs, and Transformers) regard the text as a linear sequence from the left to the right or from the right to the left. However, we humans do not always understand a document by sequentially reading it. We tend to connect the concepts across the whole document and then form structural knowledge in our brains. Motivated by the intuition, this work proposes to introduce the non-sequential connections within a document, and uses the relations among the entity mentions as the surrogates for the non-sequential conceptual connections to further improve the performance of the IE systems. Firstly, I propose to connect the related entity mentions in a document and enforce information flow among them for consistent entity type predictions. Specifically, the work connects both the local dependency relations and global coreference relations for the entity mentions to build better entity mention representations. Experimental results show that applying Graph Neural Networks (GNNs) on the connections can improve the NER performance over strong baselines on two domain-specific datasets. Secondly, I propose to connect the conceptually related regions in a document and encourage semantic interactions within and among regions. In particular, it builds the connections among the candidate role fillers (i.e., entity mentions from an event mention) for the event extraction task, and characterizes the connections based on different regional affiliations. Then edge-aware GNNs are applied to update the representations of the candidates for false positive filtering. Empirical results show that the proposed method can yield new state-of-the-art performance on two document-level event extraction datasets in two different languages. Lastly, I propose to connect the entity mentions from semi-structured tables and incorporate the table structures into the table element representations, benefiting table annotation (knowledge extraction from tables) tasks. Specifically, the system will first build hyper-graphs for the cell values coming from the same row or column and then apply the hyper-graph Neural Networks to learn better table representations. The evaluation and analysis show the effectiveness of the structure-aware table representations in improving the table annotation tasks.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor
Texas A&M University
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chen, Pei
Advisor dc:contributor.advisor
  • Huang, Ruihong
Committee members dc:contributor.committeemember
  • Ji, Shuiwang
  • Chaspari, Theodora
  • Vaid, Jyotsna

Subjects

dc:subject × 1

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1969.1/1588860

Chain of custody

source
Harvested from
Texas A&M University
Base URL
oaktrust.library.tamu.edu/server/oai/request
Last updated
2026-08-21
Source record
OAI-PMH GetRecord
citation

Chen, Pei. Connecting the Dots: Improving Information Extraction by Modeling the Non-Sequential Dependencies of Entity Mentions. Doctoral thesis, Texas A&M University, 2024. https://hdl.handle.net/1969.1/1588860