Back to results

Western Kentucky University

Efficient Schema Extraction from a Collection of XML Documents

Abstract

dc:description.abstract

<p>The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied using frequencies of the sample inputs to improve the precision of the schema extraction. Furthermore, we propose an evaluation framework to quantify the quality of the extracted schema. Experimental studies are conducted on various data sets to demonstrate the efficiency and efficacy of our approach.</p>

Degree

thesis:*
Name thesis:degree_name
Master of Science
Discipline thesis:degree_discipline
Department of Mathematics and Computer Science
Year
2011

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Parthepan, Vijayeandra
Contributors dc:contributor
  • Dr. Guangming Xing (Direcotor), Dr. Qi Li, Dr. Zhonghang Xia

Subjects

dc:subject × 5

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.wku.edu/theses/1061
OAI identifier oai:identifier
oai:digitalcommons.wku.edu:theses-2064

Chain of custody

source
Harvested from
Western Kentucky University
Base URL
digitalcommons.wku.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Parthepan, Vijayeandra. Efficient Schema Extraction from a Collection of XML Documents. 2011. https://digitalcommons.wku.edu/theses/1061