Back to results

Brigham Young University - Provo

Automating the Extraction of Domain-Specific Information from the Web-A Case Study for the Genealogical Domain

Abstract

dc:description.abstract

Current ways of finding genealogical information within the millions of pages on the Web are inadequate. In an effort to help genealogical researchers find desired information more quickly, we have developed GeneTIQS, a Genealogy Target-based Information Query System. GeneTIQS builds on ontology-based methods of data extraction to allow database-style queries on the Web. This thesis makes two main contributions to GeneTIQS. (1) It builds a framework to do generic ontology-based data extraction. (2) It develops a hybrid record separator based on Vector Space Modeling that uses both formatting clues and data clues to split pages into component records. The record separator allows GeneTIQS to extract data from the complex documents common in genealogy. Experiments show that this approach yields 92% recall and 93% precision on documents from the Web.

Degree

thesis:*
Name thesis:degree_name
MS
Grantor dc:publisher
Brigham Young University - Provo

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Walker, Troy L.

Subjects

dc:subject × 3

Rights

Language dc:language
English

Identifiers

dc:identifier.*
Repository record dc:identifier
https://scholarsarchive.byu.edu/etd/214
OAI identifier oai:identifier
oai:scholarsarchive.byu.edu:etd-1213

Chain of custody

source
Harvested from
Brigham Young University
Base URL
scholarsarchive.byu.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Walker, Troy L.. Automating the Extraction of Domain-Specific Information from the Web-A Case Study for the Genealogical Domain. Brigham Young University - Provo, https://scholarsarchive.byu.edu/etd/214