{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/16180"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/16180","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Discovery driven analysis on semi-structured text data","abstract":"Made available in DSpace on 2010-05-19T18:39:57Z (GMT). No. of bitstreams: 4 Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) 1_Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5) 2_Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) license.txt: 4064 bytes, checksum: 585cd0f4476dc2842cbbd947ebbe3b60 (MD5)","abstract_html":"Made available in DSpace on 2010-05-19T18:39:57Z (GMT). No. of bitstreams: 4 Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) 1_Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5) 2_Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) license.txt: 4064 bytes, checksum: 585cd0f4476dc2842cbbd947ebbe3b60 (MD5)","abstract_has_math":false,"creators":["Hauguel, Samson A."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-05-19T18:39:57Z","date_published":"2010-05-19T18:39:57Z","updated_at":"2026-07-22T22:25:08Z","subjects":["computer science","Information Science","Data Mining","Text Mining","Online analytical processing (OLAP)","discovery driven analysis","Probabilistic latent semantic analysis (PLSA)"],"languages":["en"],"rights":["Copyright 2010 Samson A. Hauguel"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/16180","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Hauguel, Samson A."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-05-19T18:39:57Z","2010-5"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science","Information Science","Data Mining","Text Mining","Online analytical processing (OLAP)","discovery driven analysis","Probabilistic latent semantic analysis (PLSA)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2010 Samson A. Hauguel"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/16180"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Made available in DSpace on 2010-05-19T18:39:57Z (GMT). No. of bitstreams: 4 Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) 1_Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5) 2_Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) license.txt: 4064 bytes, checksum: 585cd0f4476dc2842cbbd947ebbe3b60 (MD5)","Discovery Driven Analysis (DDA) is a common feature of OLAP technology to analyze structured data. In essence, DDA helps analysts to discover anomalous data by highlighting 'unexpected' values in the OLAP cube. By giving indications to the analyst on what dimensions to explore, DDA speeds up the process of discovering anomalies and their causes. However, Discovery Driven Analysis (and OLAP in general) is only applicable on structured data, such as records in databases. We propose a system to extend DDA technology to semi-structured text documents, that is, text documents with a few structured data. Our system pipeline consists of two stages: first, the text part of each document is structured around user specified dimensions, using semi-PLSA algorithm; then, we adapt DDA to these fully structured documents, thus enabling DDA on text documents. We present some applications of this system in OLAP analysis and show how scalability issues are solved. Results show that our system can handle reasonable datasets of documents, in real time, without any need for pre-computation.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-04-25T16:53:24Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5)"]},{"key":"dc:title","label":"Title","values":["Discovery driven analysis on semi-structured text data"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Hauguel, Samson A."],"dc:date":["2010-05-19T18:39:57Z","2010-5"],"dc:description":["Made available in DSpace on 2010-05-19T18:39:57Z (GMT). No. of bitstreams: 4 Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) 1_Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5) 2_Hauguel_Samson.pdf: 949045 bytes, checksum: c81da5764eec7003739651851a0439a5 (MD5) license.txt: 4064 bytes, checksum: 585cd0f4476dc2842cbbd947ebbe3b60 (MD5)","Discovery Driven Analysis (DDA) is a common feature of OLAP technology to analyze structured data. In essence, DDA helps analysts to discover anomalous data by highlighting 'unexpected' values in the OLAP cube. By giving indications to the analyst on what dimensions to explore, DDA speeds up the process of discovering anomalies and their causes. However, Discovery Driven Analysis (and OLAP in general) is only applicable on structured data, such as records in databases. We propose a system to extend DDA technology to semi-structured text documents, that is, text documents with a few structured data. Our system pipeline consists of two stages: first, the text part of each document is structured around user specified dimensions, using semi-PLSA algorithm; then, we adapt DDA to these fully structured documents, thus enabling DDA on text documents. We present some applications of this system in OLAP analysis and show how scalability issues are solved. Results show that our system can handle reasonable datasets of documents, in real time, without any need for pre-computation.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-04-25T16:53:24Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hauguel_Samson.pdf: 949482 bytes, checksum: 64fae416e30bc67f9ab29dac3c487ebd (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/16180"],"dc:language":["en"],"dc:rights":["Copyright 2010 Samson A. Hauguel"],"dc:subject":["computer science","Information Science","Data Mining","Text Mining","Online analytical processing (OLAP)","discovery driven analysis","Probabilistic latent semantic analysis (PLSA)"],"dc:title":["Discovery driven analysis on semi-structured text data"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:08Z"}