{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/90938"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/90938","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"TextDive: construction, summarization and exploration of multi-dimensional text corpora","abstract":"With massive datasets accumulating in text repositories (e.g., news articles, customer reviews, etc.), it is highly desirable to systematically utilize and explore them by data mining, NLP and database techniques. In our view, documents in text corpora contain informative explicit meta-attributes (e.g., category, date, author, etc.) and implicit attributes (e.g., sentiment), forming one or a set of highly-structured multi-dimensional spaces. Much knowledge can be derived if we develop effective and efficient multi-dimensional summarization, exploration and analysis technologies. In this demo, we propose an end-to-end, real-time analytical platform TextDive for processing massive text data, and provide valuable insights to general data consumers. First, we develop a set of information extraction, entity typing and text mining methods to extract consolidated dimensions and automatically construct multi-dimensional textual spaces (i.e., text cubes). Furthermore, we develop a set of OLAP-like text summarization, data exploration and text analysis mechanisms that understand semantics of text corpora in multi-dimensional spaces. We also develop an efficient computational solution that involves materializing selective statistics to guarantee the interactive and real-time nature of TextDive.","abstract_html":"With massive datasets accumulating in text repositories (e.g., news articles, customer reviews, etc.), it is highly desirable to systematically utilize and explore them by data mining, NLP and database techniques. In our view, documents in text corpora contain informative explicit meta-attributes (e.g., category, date, author, etc.) and implicit attributes (e.g., sentiment), forming one or a set of highly-structured multi-dimensional spaces. Much knowledge can be derived if we develop effective and efficient multi-dimensional summarization, exploration and analysis technologies. In this demo, we propose an end-to-end, real-time analytical platform TextDive for processing massive text data, and provide valuable insights to general data consumers. First, we develop a set of information extraction, entity typing and text mining methods to extract consolidated dimensions and automatically construct multi-dimensional textual spaces (i.e., text cubes). Furthermore, we develop a set of OLAP-like text summarization, data exploration and text analysis mechanisms that understand semantics of text corpora in multi-dimensional spaces. We also develop an efficient computational solution that involves materializing selective statistics to guarantee the interactive and real-time nature of TextDive.","abstract_has_math":false,"creators":["Wang, Qi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-07-07T21:17:51Z","date_published":"2016-07-07T21:17:51Z","updated_at":"2026-07-22T22:26:34Z","subjects":["multi-dimensional text corpora analysis","text cube analysis","text summarization"],"languages":["en"],"rights":["Copyright 2016 Qi Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/90938","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Wang, Qi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-07-07T21:17:51Z","2018-07-08T09:15:16Z","2016-04-20","2016-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["multi-dimensional text corpora analysis","text cube analysis","text summarization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Qi Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/90938"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["With massive datasets accumulating in text repositories (e.g., news articles, customer reviews, etc.), it is highly desirable to systematically utilize and explore them by data mining, NLP and database techniques. In our view, documents in text corpora contain informative explicit meta-attributes (e.g., category, date, author, etc.) and implicit attributes (e.g., sentiment), forming one or a set of highly-structured multi-dimensional spaces. Much knowledge can be derived if we develop effective and efficient multi-dimensional summarization, exploration and analysis technologies. In this demo, we propose an end-to-end, real-time analytical platform TextDive for processing massive text data, and provide valuable insights to general data consumers. First, we develop a set of information extraction, entity typing and text mining methods to extract consolidated dimensions and automatically construct multi-dimensional textual spaces (i.e., text cubes). Furthermore, we develop a set of OLAP-like text summarization, data exploration and text analysis mechanisms that understand semantics of text corpora in multi-dimensional spaces. We also develop an efficient computational solution that involves materializing selective statistics to guarantee the interactive and real-time nature of TextDive.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-05-01","The student, Qi Wang, accepted the attached license on 2016-04-20 at 14:36.","The student, Qi Wang, submitted this Thesis for approval on 2016-04-20 at 14:52.","This Thesis was approved for publication on 2016-04-20 at 16:51.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9379 on 2016-07-07 at 14:17:35","Made available in DSpace on 2016-07-07T21:17:51Z (GMT). No. of bitstreams: 2 WANG-THESIS-2016.pdf: 741348 bytes, checksum: 5341f44b7d61b90abaa498309d081efb (MD5) LICENSE.txt: 4204 bytes, checksum: e8e3437e2af1b48d05e9678958b4f830 (MD5) Previous issue date: 2016-04-20","Embargo set by: Seth Robbins for item 93293 Lift date: 2018-07-07T21:18:16Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 93293 on 2018-07-08T09:15:16Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["TextDive: construction, summarization and exploration of multi-dimensional text corpora"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Wang, Qi"],"dc:date":["2016-07-07T21:17:51Z","2018-07-08T09:15:16Z","2016-04-20","2016-05"],"dc:description":["With massive datasets accumulating in text repositories (e.g., news articles, customer reviews, etc.), it is highly desirable to systematically utilize and explore them by data mining, NLP and database techniques. In our view, documents in text corpora contain informative explicit meta-attributes (e.g., category, date, author, etc.) and implicit attributes (e.g., sentiment), forming one or a set of highly-structured multi-dimensional spaces. Much knowledge can be derived if we develop effective and efficient multi-dimensional summarization, exploration and analysis technologies. In this demo, we propose an end-to-end, real-time analytical platform TextDive for processing massive text data, and provide valuable insights to general data consumers. First, we develop a set of information extraction, entity typing and text mining methods to extract consolidated dimensions and automatically construct multi-dimensional textual spaces (i.e., text cubes). Furthermore, we develop a set of OLAP-like text summarization, data exploration and text analysis mechanisms that understand semantics of text corpora in multi-dimensional spaces. We also develop an efficient computational solution that involves materializing selective statistics to guarantee the interactive and real-time nature of TextDive.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-05-01","The student, Qi Wang, accepted the attached license on 2016-04-20 at 14:36.","The student, Qi Wang, submitted this Thesis for approval on 2016-04-20 at 14:52.","This Thesis was approved for publication on 2016-04-20 at 16:51.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9379 on 2016-07-07 at 14:17:35","Made available in DSpace on 2016-07-07T21:17:51Z (GMT). No. of bitstreams: 2 WANG-THESIS-2016.pdf: 741348 bytes, checksum: 5341f44b7d61b90abaa498309d081efb (MD5) LICENSE.txt: 4204 bytes, checksum: e8e3437e2af1b48d05e9678958b4f830 (MD5) Previous issue date: 2016-04-20","Embargo set by: Seth Robbins for item 93293 Lift date: 2018-07-07T21:18:16Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 93293 on 2018-07-08T09:15:16Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/90938"],"dc:language":["en"],"dc:rights":["Copyright 2016 Qi Wang"],"dc:subject":["multi-dimensional text corpora analysis","text cube analysis","text summarization"],"dc:title":["TextDive: construction, summarization and exploration of multi-dimensional text corpora"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:34Z"}