{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81869"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81869","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large-Scale Constraint-Based Pattern Mining","abstract":"We studied the problem of constraint-based pattern mining for three different data formats, item-set, sequence and graph, and focused on mining patterns of large sizes. Colossal patterns in each data formats are studied to discover pruning properties that are useful for direct mining of these patterns. For item-set data, we observed robustness of colossal patterns. By defining the concept of core patterns, we developed a randomized mining framework to efficiently find the set of colossal patterns which gives a good approximation to the complete pattern set. The essential idea of pattern fusion and leaping toward large patterns is then extended to the cases of sequential and graph data. In sequential data, we developed a novel algorithm to accommodate approximate patterns. For graph data, we proposed the concept of spiders and used these pre-computed frequent structures of small sizes to quickly leap to reach those much larger ones. We also proposed a general graph mining framework, called gPrune, to take advantage of both pattern and data space pruning. Ideas and techniques developed in this work can be extended to handle other user-specified constraints for direct efficient mining in large-scale data.","abstract_html":"We studied the problem of constraint-based pattern mining for three different data formats, item-set, sequence and graph, and focused on mining patterns of large sizes. Colossal patterns in each data formats are studied to discover pruning properties that are useful for direct mining of these patterns. For item-set data, we observed robustness of colossal patterns. By defining the concept of core patterns, we developed a randomized mining framework to efficiently find the set of colossal patterns which gives a good approximation to the complete pattern set. The essential idea of pattern fusion and leaping toward large patterns is then extended to the cases of sequential and graph data. In sequential data, we developed a novel algorithm to accommodate approximate patterns. For graph data, we proposed the concept of spiders and used these pre-computed frequent structures of small sizes to quickly leap to reach those much larger ones. We also proposed a general graph mining framework, called gPrune, to take advantage of both pattern and data space pruning. Ideas and techniques developed in this work can be extended to handle other user-specified constraints for direct efficient mining in large-scale data.","abstract_has_math":false,"creators":["Zhu, Feida"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Jeff Erickson","Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:20:46Z","date_published":"2015-09-25T20:20:46Z","updated_at":"2026-07-22T22:26:17Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3395564"],"render_values":[{"text":"(MiAaPQ)AAI3395564","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81869","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jeff Erickson","Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Zhu, Feida"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:20:46Z","10000-01-01","2009"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81869","(MiAaPQ)AAI3395564"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We studied the problem of constraint-based pattern mining for three different data formats, item-set, sequence and graph, and focused on mining patterns of large sizes. Colossal patterns in each data formats are studied to discover pruning properties that are useful for direct mining of these patterns. For item-set data, we observed robustness of colossal patterns. By defining the concept of core patterns, we developed a randomized mining framework to efficiently find the set of colossal patterns which gives a good approximation to the complete pattern set. The essential idea of pattern fusion and leaping toward large patterns is then extended to the cases of sequential and graph data. In sequential data, we developed a novel algorithm to accommodate approximate patterns. For graph data, we proposed the concept of spiders and used these pre-computed frequent structures of small sizes to quickly leap to reach those much larger ones. We also proposed a general graph mining framework, called gPrune, to take advantage of both pattern and data space pruning. Ideas and techniques developed in this work can be extended to handle other user-specified constraints for direct efficient mining in large-scale data.","Made available in DSpace on 2015-09-25T20:20:46Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3395564.pdf: 1447460 bytes, checksum: 9b556e8c4f19efe395b13a12a2874326 (MD5) Previous issue date: 2009","Embargo set by: Seth Robbins for item 83150 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","89 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2009."]},{"key":"dc:title","label":"Title","values":["Large-Scale Constraint-Based Pattern Mining"]}]}],"canonical_facts":{"dc:contributor":["Jeff Erickson","Han, Jiawei"],"dc:creator":["Zhu, Feida"],"dc:date":["2015-09-25T20:20:46Z","10000-01-01","2009"],"dc:description":["We studied the problem of constraint-based pattern mining for three different data formats, item-set, sequence and graph, and focused on mining patterns of large sizes. Colossal patterns in each data formats are studied to discover pruning properties that are useful for direct mining of these patterns. For item-set data, we observed robustness of colossal patterns. By defining the concept of core patterns, we developed a randomized mining framework to efficiently find the set of colossal patterns which gives a good approximation to the complete pattern set. The essential idea of pattern fusion and leaping toward large patterns is then extended to the cases of sequential and graph data. In sequential data, we developed a novel algorithm to accommodate approximate patterns. For graph data, we proposed the concept of spiders and used these pre-computed frequent structures of small sizes to quickly leap to reach those much larger ones. We also proposed a general graph mining framework, called gPrune, to take advantage of both pattern and data space pruning. Ideas and techniques developed in this work can be extended to handle other user-specified constraints for direct efficient mining in large-scale data.","Made available in DSpace on 2015-09-25T20:20:46Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3395564.pdf: 1447460 bytes, checksum: 9b556e8c4f19efe395b13a12a2874326 (MD5) Previous issue date: 2009","Embargo set by: Seth Robbins for item 83150 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","89 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2009."],"dc:identifier":["http://hdl.handle.net/2142/81869","(MiAaPQ)AAI3395564"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Large-Scale Constraint-Based Pattern Mining"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:17Z"}