{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104764"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104764","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Extracting and utilizing hidden structures in large datasets","abstract":"The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to potentially surpass traditional algorithms that do not take this structure into account. In this thesis, we propose a general framework for algorithms that automatically extract and employ hidden structures to improve data processing performance, and discuss a set of design principles for developing such algorithms. We provide three examples to demonstrate the power of this framework in practice, showcasing how we can use hidden structures to either outperform state-of-the-art methods, or enable new applications that are previously impossible. We believe that this framework can offer new opportunities for the design of algorithms that surpass the current limit, and empower new applications in database research and many other data-centric disciplines.","abstract_html":"The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to potentially surpass traditional algorithms that do not take this structure into account. In this thesis, we propose a general framework for algorithms that automatically extract and employ hidden structures to improve data processing performance, and discuss a set of design principles for developing such algorithms. We provide three examples to demonstrate the power of this framework in practice, showcasing how we can use hidden structures to either outperform state-of-the-art methods, or enable new applications that are previously impossible. We believe that this framework can offer new opportunities for the design of algorithms that surpass the current limit, and empower new applications in database research and many other data-centric disciplines.","abstract_has_math":false,"creators":["Gao, Yihan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Parameswaran, Aditya","Chang, Kevin","Sundaram, Hari","Wang, Jiannan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T19:51:33Z","date_published":"2019-08-23T19:51:33Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Structure Extraction","Automatic Data Processing"],"languages":["en"],"rights":["Copyright 2019 Yihan Gao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104764","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Parameswaran, Aditya","Chang, Kevin","Sundaram, Hari","Wang, Jiannan"]},{"key":"dc:creator","label":"Author","values":["Gao, Yihan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T19:51:33Z","2019-03-26","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Structure Extraction","Automatic Data Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yihan Gao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104764"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to potentially surpass traditional algorithms that do not take this structure into account. In this thesis, we propose a general framework for algorithms that automatically extract and employ hidden structures to improve data processing performance, and discuss a set of design principles for developing such algorithms. We provide three examples to demonstrate the power of this framework in practice, showcasing how we can use hidden structures to either outperform state-of-the-art methods, or enable new applications that are previously impossible. We believe that this framework can offer new opportunities for the design of algorithms that surpass the current limit, and empower new applications in database research and many other data-centric disciplines.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yihan Gao, accepted the attached license on 2019-03-23 at 02:47.","The student, Yihan Gao, submitted this Dissertation for approval on 2019-03-23 at 02:58.","This Dissertation was approved for publication on 2019-03-26 at 18:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13445 on 2019-08-22 at 14:40:44","Made available in DSpace on 2019-08-23T19:51:33Z (GMT). No. of bitstreams: 2 GAO-DISSERTATION-2019.pdf: 5453832 bytes, checksum: 1d959e6473c77a5f1f53d2b3a919ff7d (MD5) LICENSE.txt: 4206 bytes, checksum: 5855723f629b84003cf5854c0b850f4d (MD5) Previous issue date: 2019-03-26"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Extracting and utilizing hidden structures in large datasets"]}]}],"canonical_facts":{"dc:contributor":["Parameswaran, Aditya","Chang, Kevin","Sundaram, Hari","Wang, Jiannan"],"dc:creator":["Gao, Yihan"],"dc:date":["2019-08-23T19:51:33Z","2019-03-26","2019-05"],"dc:description":["The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to potentially surpass traditional algorithms that do not take this structure into account. In this thesis, we propose a general framework for algorithms that automatically extract and employ hidden structures to improve data processing performance, and discuss a set of design principles for developing such algorithms. We provide three examples to demonstrate the power of this framework in practice, showcasing how we can use hidden structures to either outperform state-of-the-art methods, or enable new applications that are previously impossible. We believe that this framework can offer new opportunities for the design of algorithms that surpass the current limit, and empower new applications in database research and many other data-centric disciplines.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yihan Gao, accepted the attached license on 2019-03-23 at 02:47.","The student, Yihan Gao, submitted this Dissertation for approval on 2019-03-23 at 02:58.","This Dissertation was approved for publication on 2019-03-26 at 18:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13445 on 2019-08-22 at 14:40:44","Made available in DSpace on 2019-08-23T19:51:33Z (GMT). No. of bitstreams: 2 GAO-DISSERTATION-2019.pdf: 5453832 bytes, checksum: 1d959e6473c77a5f1f53d2b3a919ff7d (MD5) LICENSE.txt: 4206 bytes, checksum: 5855723f629b84003cf5854c0b850f4d (MD5) Previous issue date: 2019-03-26"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104764"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yihan Gao"],"dc:subject":["Structure Extraction","Automatic Data Processing"],"dc:title":["Extracting and utilizing hidden structures in large datasets"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}