Back to results

University of New Orleans

A Semi-Supervised Information Extraction Framework for Large Redundant Corpora

Abstract

dc:description.abstract

The vast majority of text freely available on the Internet is not available in a form that computers can understand. There have been numerous approaches to automatically extract information from human- readable sources. The most successful attempts rely on vast training sets of data. Others have succeeded in extracting restricted subsets of the available information. These approaches have limited use and require domain knowledge to be coded into the application. The current thesis proposes a novel framework for Information Extraction. From large sets of documents, the system develops statistical models of the data the user wishes to query which generally avoid the lim- itations and complexity of most Information Extractions systems. The framework uses a semi-supervised approach to minimize human input. It also eliminates the need for external Named Entity Recognition systems by relying on freely available databases. The final result is a query-answering system which extracts information from large corpora with a high degree of accuracy.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Year
2008

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Normand, Eric
Contributors dc:contributor
  • Abdelguerfi, Mahdi
  • Richard III, Golden
  • Tu, Shengru

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Repository record dc:identifier
https://scholarworks.uno.edu/td/877
OAI identifier oai:identifier
oai:scholarworks.uno.edu:td-1857

Chain of custody

source
Harvested from
University of New Orleans
Base URL
scholarworks.uno.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Normand, Eric. A Semi-Supervised Information Extraction Framework for Large Redundant Corpora. Thesis thesis, 2008. https://scholarworks.uno.edu/td/877