{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/86731"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/86731","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Summarizing Semi-Structured Data","abstract":"Ph.D.","abstract_html":"Ph.D.","abstract_has_math":false,"creators":["Xie, Ting; 0000-0002-3092-9975"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Kennedy, Oliver","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-02-21T21:44:31Z","date_published":"2025-02-21T21:44:31Z","updated_at":"2026-07-27T19:05:34Z","subjects":["computer science"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/86731","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kennedy, Oliver","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Xie, Ting; 0000-0002-3092-9975"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-02-21T21:44:31Z","2020"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/86731"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Ph.D.","Data sources, e.g., Yelp or Twitter, that produce records without pre-defined schema (semi-structured) are popular nowadays. However, semi-structured data can be unwieldy for data exploration purposes due to an extremely large number of records as well as an accumulative number of attributes. Finding a more compact data representation (i.e., summary encoding) is thus critical before subsequent data analysis can be applied. Designing an appropriate summary encoding requires a trade-off between compactness and fidelity. In this thesis, a framework is developed for reasoning about the trade-off between summary compactness and fidelity. A measure of summary fidelity is proposed accordingly. Analytic and experimental evidences have been presented, which show that the proposed fidelity measure is not only efficiently computable, but also a meaningful measure of summary quality. To efficiently construct a high-fidelity summary encoding given a limitation on its verboseness, a clustering-based approach is identified and experiment results show that it produces results orders of magnitude faster and competitive with more powerful techniques for compression and summarization.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Summarizing Semi-Structured Data"]}]}],"canonical_facts":{"dc:contributor":["Kennedy, Oliver","Computer Science and Engineering"],"dc:creator":["Xie, Ting; 0000-0002-3092-9975"],"dc:date":["2025-02-21T21:44:31Z","2020"],"dc:description":["Ph.D.","Data sources, e.g., Yelp or Twitter, that produce records without pre-defined schema (semi-structured) are popular nowadays. However, semi-structured data can be unwieldy for data exploration purposes due to an extremely large number of records as well as an accumulative number of attributes. Finding a more compact data representation (i.e., summary encoding) is thus critical before subsequent data analysis can be applied. Designing an appropriate summary encoding requires a trade-off between compactness and fidelity. In this thesis, a framework is developed for reasoning about the trade-off between summary compactness and fidelity. A measure of summary fidelity is proposed accordingly. Analytic and experimental evidences have been presented, which show that the proposed fidelity measure is not only efficiently computable, but also a meaningful measure of summary quality. To efficiently construct a high-fidelity summary encoding given a limitation on its verboseness, a clustering-based approach is identified and experiment results show that it produces results orders of magnitude faster and competitive with more powerful techniques for compression and summarization.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/86731"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science"],"dc:title":["Summarizing Semi-Structured Data"],"dc:type":["Text","Dissertation"]},"updated_at":"2026-07-27T19:05:34Z"}