{"id":{"repo_id":"wku-diss","oai_identifier":"oai:digitalcommons.wku.edu:theses-2064"},"canonical_url":"https://search.dev.ndltd.org/etd/wku-diss/oai:digitalcommons.wku.edu:theses-2064","repository":{"repo_id":"wku-diss","name":"Western Kentucky University","base_url":"https://digitalcommons.wku.edu/do/oai/"},"display":{"title":"Efficient Schema Extraction from a Collection of XML Documents","abstract":"<p>The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied using frequencies of the sample inputs to improve the precision of the schema extraction. Furthermore, we propose an evaluation framework to quantify the quality of the extracted schema. Experimental studies are conducted on various data sets to demonstrate the efficiency and efficacy of our approach.</p>","abstract_html":"&lt;p&gt;The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied using frequencies of the sample inputs to improve the precision of the schema extraction. Furthermore, we propose an evaluation framework to quantify the quality of the extracted schema. Experimental studies are conducted on various data sets to demonstrate the efficiency and efficacy of our approach.&lt;/p&gt;","abstract_has_math":false,"creators":["Parthepan, Vijayeandra"],"institution":null,"degree_name":"Master of Science","degree_level":null,"degree_discipline":"Department of Mathematics and Computer Science","degree_department":null,"school":null,"contributors":["Dr. Guangming Xing (Direcotor), Dr. Qi Li, Dr. Zhonghang Xia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-01T07:00:00Z","date_published":"2011-05-01T07:00:00Z","updated_at":"2026-07-24T06:08:09Z","subjects":["document mark up language","data mining","eXtensible Markup Language","Databases and Information Systems","Programming Languages and Compilers"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.wku.edu/theses/1061","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Guangming Xing (Direcotor), Dr. Qi Li, Dr. Zhonghang Xia"]},{"key":"dc:creator","label":"Author","values":["Parthepan, Vijayeandra"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Department of Mathematics and Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["document mark up language","data mining","eXtensible Markup Language","Databases and Information Systems","Programming Languages and Compilers"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.wku.edu/theses/1061"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied using frequencies of the sample inputs to improve the precision of the schema extraction. Furthermore, we propose an evaluation framework to quantify the quality of the extracted schema. Experimental studies are conducted on various data sets to demonstrate the efficiency and efficacy of our approach.</p>"]},{"key":"dc:title","label":"Title","values":["Efficient Schema Extraction from a Collection of XML Documents"]}]}],"canonical_facts":{"dc:contributor":["Dr. Guangming Xing (Direcotor), Dr. Qi Li, Dr. Zhonghang Xia"],"dc:creator":["Parthepan, Vijayeandra"],"dc:description.abstract":["<p>The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied using frequencies of the sample inputs to improve the precision of the schema extraction. Furthermore, we propose an evaluation framework to quantify the quality of the extracted schema. Experimental studies are conducted on various data sets to demonstrate the efficiency and efficacy of our approach.</p>"],"dc:identifier":["https://digitalcommons.wku.edu/theses/1061"],"dc:subject":["document mark up language","data mining","eXtensible Markup Language","Databases and Information Systems","Programming Languages and Compilers"],"dc:title":["Efficient Schema Extraction from a Collection of XML Documents"],"dc:type":["Thesis"],"thesis:degree_discipline":["Department of Mathematics and Computer Science"],"thesis:degree_name":["Master of Science"]},"updated_at":"2026-07-24T06:08:09Z"}