Back to results

Western Kentucky University

Automatically Extract Information from Web Documents

Abstract

dc:description.abstract

The Internet could be considered to be a reservoir of useful information in textual form — product catalogs, airline schedules, stock market quotations, weather forecast etc. There has been much interest in building systems that gather such information on a user's behalf. But because these information resources are formatted differently, mechanically extracting their content is difficult. Systems using such resources typically use hand-coded wrappers, customized procedures for information extraction. Structured data objects are a very important type of information on the Web. Such data objects are often records from underlying databases and displayed in Web pages with some fixed templates. Mining data records in Web pages is useful because they typically present their host pages' essential information, such as lists of products and services. Extracting these structured data objects enables one to integrate data/information from multiple Web pages to provide value-added services, e.g., comparative shopping, meta-querying and search. Web content mining has thus become an area of interest for many researchers because of the phenomenal growth of the Web contents and the economic benefits associated with it. However, due to the heterogeneity of Web pages, automated discovery of targeted information is still posing as a challenging problem.

Degree

thesis:*
Name thesis:degree_name
Master of Science in Computer Science
Discipline thesis:degree_discipline
Department of Mathematics and Computer Science
Year
2007

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sharma, Dipesh

Subjects

dc:subject × 2

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.wku.edu/theses/376
OAI identifier oai:identifier
oai:digitalcommons.wku.edu:theses-1379

Chain of custody

source
Harvested from
Western Kentucky University
Base URL
digitalcommons.wku.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Sharma, Dipesh. Automatically Extract Information from Web Documents. 2007. https://digitalcommons.wku.edu/theses/376