{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/99244"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/99244","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Text cube: construction, summarization and mining","abstract":"A large portion of real world data is either text or structured (\\eg, relational) data. Such data objects are often linked together (\\eg, structured product information linking with their descriptions and customer reviews.). To systematically analyze large numbers of such textual documents, it is often desirable to manage the text data with the associated structured data in a multi-dimensional space (hence \\emph{text cube}). This thesis studies the multi-dimensional representation of large textual data. Since Jim Gray introduced the concept of ``data cube'', data cube, associated with online analytical processing (OLAP), has become a driving engine in data warehouse industry. By modeling a large textual corpus as a ``cube'', \\ie multi-dimensional and hierarchical structure, we bridge the power of traditional OLAP and Information Retrieval / Natural Language Processing techniques. In particular, this thesis focuses on two lines of work, one is to construct a multi-dimensional text cube from raw text data with limited user guidance; the other is to develop effective summarization and mining techniques tailored for multi-dimensional queries on text cubes. In the first part of the thesis, the problem of \\emph{dimension-based structure creation} is studied. We propose an end-to-end framework for extracting multi-dimensional structure from a corpus, taking the input of a corpus of specific domain and limited seeds to generate a high-quality dimension values as output. We introduce the novel concept of Semantic Pattern Graph to leverage web signals to understand the underlying semantics of lexical patterns, improve pattern evaluation using mined semantics, and yield more accurate and complete structure. Experiments show the effectiveness of our approach. In the second part, with all the dimensions discovered, we study the problem of \\emph{cell-based document allocation}. That is, linking the created dimensions with text data and construct a multi-dimensional text cube. To allocate documents into correct multi-dimensional subsets, \\ie a cell. Traditional approaches, in this particular task, may require substantial labeling from user. Instead, we propose a model that requires no additional training data besides the given (label) name of each cube dimension as weak supervision. With such weak supervision, we develop a \\emph{dimension-aware joint embedding} framework that learns joint representations for terms, documents, and labels. In the joint embedding process, our method iteratively learns dimension-aware document representations by selectively focusing on discriminative keywords for different dimensions. Furthermore, it alleviates label sparsity by leveraging label representations to enrich the labeled term set. Numerical experiments corroborate the effectiveness of our solution. In the third part, we introduce the concept of \\emph{Context-Aware Semantic Online Analytical Processing} (\\ie \\emph{CASeOLAP}) in text cubes, and use \\emph{top-$k$ representative phrases} to represent the semantics of the document subset in a text cube cell. By ranking phrases with a newly proposed ranking measure according to three criteria: integrity, popularity and distinctiveness. We identify phrases that can successfully digest the main content of a subset of documents of interest and contrast with other neighboring subsets. Our experiments in a large news dataset demonstrate the effectiveness of the newly proposed ranking measure in finding representative phrases and the efficiency in both query processing time and storage cost. The approach is also applied to clinical biomarker analysis and protest news analysis with success. In the last part, the system of \\emph{EventCube} is proposed to support end-to-end pipeline of text cube in an informative, interactive, and user-friendly manner. The system serves as a general platform for construction, search, summarization, OLAP (online analytical processing) and data mining on integrated text and structured data. The system is a growing testbed for various text cube based research and has been successfully applied to NASA for aviation safety report analysis and Army Research Lab for Counter-Terrorism Report analysis. To summarize, this thesis provides important results of construction and consumption of multi-dimensional text cubes and shows its power in tackling real-world text analysis tasks.","abstract_html":"A large portion of real world data is either text or structured (\\eg, relational) data. Such data objects are often linked together (\\eg, structured product information linking with their descriptions and customer reviews.). To systematically analyze large numbers of such textual documents, it is often desirable to manage the text data with the associated structured data in a multi-dimensional space (hence \\emph{text cube}). This thesis studies the multi-dimensional representation of large textual data. Since Jim Gray introduced the concept of ``data cube&#x27;&#x27;, data cube, associated with online analytical processing (OLAP), has become a driving engine in data warehouse industry. By modeling a large textual corpus as a ``cube&#x27;&#x27;, \\ie multi-dimensional and hierarchical structure, we bridge the power of traditional OLAP and Information Retrieval / Natural Language Processing techniques. In particular, this thesis focuses on two lines of work, one is to construct a multi-dimensional text cube from raw text data with limited user guidance; the other is to develop effective summarization and mining techniques tailored for multi-dimensional queries on text cubes. In the first part of the thesis, the problem of \\emph{dimension-based structure creation} is studied. We propose an end-to-end framework for extracting multi-dimensional structure from a corpus, taking the input of a corpus of specific domain and limited seeds to generate a high-quality dimension values as output. We introduce the novel concept of Semantic Pattern Graph to leverage web signals to understand the underlying semantics of lexical patterns, improve pattern evaluation using mined semantics, and yield more accurate and complete structure. Experiments show the effectiveness of our approach. In the second part, with all the dimensions discovered, we study the problem of \\emph{cell-based document allocation}. That is, linking the created dimensions with text data and construct a multi-dimensional text cube. To allocate documents into correct multi-dimensional subsets, \\ie a cell. Traditional approaches, in this particular task, may require substantial labeling from user. Instead, we propose a model that requires no additional training data besides the given (label) name of each cube dimension as weak supervision. With such weak supervision, we develop a \\emph{dimension-aware joint embedding} framework that learns joint representations for terms, documents, and labels. In the joint embedding process, our method iteratively learns dimension-aware document representations by selectively focusing on discriminative keywords for different dimensions. Furthermore, it alleviates label sparsity by leveraging label representations to enrich the labeled term set. Numerical experiments corroborate the effectiveness of our solution. In the third part, we introduce the concept of \\emph{Context-Aware Semantic Online Analytical Processing} (\\ie \\emph{CASeOLAP}) in text cubes, and use \\emph{top-$k$ representative phrases} to represent the semantics of the document subset in a text cube cell. By ranking phrases with a newly proposed ranking measure according to three criteria: integrity, popularity and distinctiveness. We identify phrases that can successfully digest the main content of a subset of documents of interest and contrast with other neighboring subsets. Our experiments in a large news dataset demonstrate the effectiveness of the newly proposed ranking measure in finding representative phrases and the efficiency in both query processing time and storage cost. The approach is also applied to clinical biomarker analysis and protest news analysis with success. In the last part, the system of \\emph{EventCube} is proposed to support end-to-end pipeline of text cube in an informative, interactive, and user-friendly manner. The system serves as a general platform for construction, search, summarization, OLAP (online analytical processing) and data mining on integrated text and structured data. The system is a growing testbed for various text cube based research and has been successfully applied to NASA for aviation safety report analysis and Army Research Lab for Counter-Terrorism Report analysis. To summarize, this thesis provides important results of construction and consumption of multi-dimensional text cubes and shows its power in tackling real-world text analysis tasks.","abstract_has_math":true,"creators":["Tao, Fangbo"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Zhai, ChengXiang","Peng, Jian","Wang, Haixun"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-03-13T15:25:31Z","date_published":"2018-03-13T15:25:31Z","updated_at":"2026-07-22T22:24:37Z","subjects":["Text cube","Data cube","Data mining","Natural language processing","Text classification"],"languages":["en"],"rights":["Copyright 2017 Fangbo Tao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/99244","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Zhai, ChengXiang","Peng, Jian","Wang, Haixun"]},{"key":"dc:creator","label":"Author","values":["Tao, Fangbo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-03-13T15:25:31Z","2020-03-14T09:15:25Z","2017-12-06","2017-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Text cube","Data cube","Data mining","Natural language processing","Text classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Fangbo Tao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/99244"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A large portion of real world data is either text or structured (\\eg, relational) data. Such data objects are often linked together (\\eg, structured product information linking with their descriptions and customer reviews.). To systematically analyze large numbers of such textual documents, it is often desirable to manage the text data with the associated structured data in a multi-dimensional space (hence \\emph{text cube}). This thesis studies the multi-dimensional representation of large textual data. Since Jim Gray introduced the concept of ``data cube'', data cube, associated with online analytical processing (OLAP), has become a driving engine in data warehouse industry. By modeling a large textual corpus as a ``cube'', \\ie multi-dimensional and hierarchical structure, we bridge the power of traditional OLAP and Information Retrieval / Natural Language Processing techniques. In particular, this thesis focuses on two lines of work, one is to construct a multi-dimensional text cube from raw text data with limited user guidance; the other is to develop effective summarization and mining techniques tailored for multi-dimensional queries on text cubes. In the first part of the thesis, the problem of \\emph{dimension-based structure creation} is studied. We propose an end-to-end framework for extracting multi-dimensional structure from a corpus, taking the input of a corpus of specific domain and limited seeds to generate a high-quality dimension values as output. We introduce the novel concept of Semantic Pattern Graph to leverage web signals to understand the underlying semantics of lexical patterns, improve pattern evaluation using mined semantics, and yield more accurate and complete structure. Experiments show the effectiveness of our approach. In the second part, with all the dimensions discovered, we study the problem of \\emph{cell-based document allocation}. That is, linking the created dimensions with text data and construct a multi-dimensional text cube. To allocate documents into correct multi-dimensional subsets, \\ie a cell. Traditional approaches, in this particular task, may require substantial labeling from user. Instead, we propose a model that requires no additional training data besides the given (label) name of each cube dimension as weak supervision. With such weak supervision, we develop a \\emph{dimension-aware joint embedding} framework that learns joint representations for terms, documents, and labels. In the joint embedding process, our method iteratively learns dimension-aware document representations by selectively focusing on discriminative keywords for different dimensions. Furthermore, it alleviates label sparsity by leveraging label representations to enrich the labeled term set. Numerical experiments corroborate the effectiveness of our solution. In the third part, we introduce the concept of \\emph{Context-Aware Semantic Online Analytical Processing} (\\ie \\emph{CASeOLAP}) in text cubes, and use \\emph{top-$k$ representative phrases} to represent the semantics of the document subset in a text cube cell. By ranking phrases with a newly proposed ranking measure according to three criteria: integrity, popularity and distinctiveness. We identify phrases that can successfully digest the main content of a subset of documents of interest and contrast with other neighboring subsets. Our experiments in a large news dataset demonstrate the effectiveness of the newly proposed ranking measure in finding representative phrases and the efficiency in both query processing time and storage cost. The approach is also applied to clinical biomarker analysis and protest news analysis with success. In the last part, the system of \\emph{EventCube} is proposed to support end-to-end pipeline of text cube in an informative, interactive, and user-friendly manner. The system serves as a general platform for construction, search, summarization, OLAP (online analytical processing) and data mining on integrated text and structured data. The system is a growing testbed for various text cube based research and has been successfully applied to NASA for aviation safety report analysis and Army Research Lab for Counter-Terrorism Report analysis. To summarize, this thesis provides important results of construction and consumption of multi-dimensional text cubes and shows its power in tackling real-world text analysis tasks.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-12-01","The student, Fangbo Tao, accepted the attached license on 2017-12-06 at 11:45.","The student, Fangbo Tao, submitted this Dissertation for approval on 2017-12-06 at 11:53.","This Dissertation was approved for publication on 2017-12-06 at 13:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11883 on 2018-03-13 at 09:57:28","Made available in DSpace on 2018-03-13T15:25:31Z (GMT). No. of bitstreams: 2 TAO-DISSERTATION-2017.pdf: 7757301 bytes, checksum: eb265911f521a46b4473b8c6145b02e0 (MD5) LICENSE.txt: 4207 bytes, checksum: 441a8dcbebe919d9cdf10375f5c5b7ae (MD5) Previous issue date: 2017-12-06","Embargo set by: Seth Robbins for item 105207 Lift date: 2020-03-13T15:25:40Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105207 Lift date: 2020-03-13T15:28:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105207 on 2020-03-14T09:15:25Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Text cube: construction, summarization and mining"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Zhai, ChengXiang","Peng, Jian","Wang, Haixun"],"dc:creator":["Tao, Fangbo"],"dc:date":["2018-03-13T15:25:31Z","2020-03-14T09:15:25Z","2017-12-06","2017-12"],"dc:description":["A large portion of real world data is either text or structured (\\eg, relational) data. Such data objects are often linked together (\\eg, structured product information linking with their descriptions and customer reviews.). To systematically analyze large numbers of such textual documents, it is often desirable to manage the text data with the associated structured data in a multi-dimensional space (hence \\emph{text cube}). This thesis studies the multi-dimensional representation of large textual data. Since Jim Gray introduced the concept of ``data cube'', data cube, associated with online analytical processing (OLAP), has become a driving engine in data warehouse industry. By modeling a large textual corpus as a ``cube'', \\ie multi-dimensional and hierarchical structure, we bridge the power of traditional OLAP and Information Retrieval / Natural Language Processing techniques. In particular, this thesis focuses on two lines of work, one is to construct a multi-dimensional text cube from raw text data with limited user guidance; the other is to develop effective summarization and mining techniques tailored for multi-dimensional queries on text cubes. In the first part of the thesis, the problem of \\emph{dimension-based structure creation} is studied. We propose an end-to-end framework for extracting multi-dimensional structure from a corpus, taking the input of a corpus of specific domain and limited seeds to generate a high-quality dimension values as output. We introduce the novel concept of Semantic Pattern Graph to leverage web signals to understand the underlying semantics of lexical patterns, improve pattern evaluation using mined semantics, and yield more accurate and complete structure. Experiments show the effectiveness of our approach. In the second part, with all the dimensions discovered, we study the problem of \\emph{cell-based document allocation}. That is, linking the created dimensions with text data and construct a multi-dimensional text cube. To allocate documents into correct multi-dimensional subsets, \\ie a cell. Traditional approaches, in this particular task, may require substantial labeling from user. Instead, we propose a model that requires no additional training data besides the given (label) name of each cube dimension as weak supervision. With such weak supervision, we develop a \\emph{dimension-aware joint embedding} framework that learns joint representations for terms, documents, and labels. In the joint embedding process, our method iteratively learns dimension-aware document representations by selectively focusing on discriminative keywords for different dimensions. Furthermore, it alleviates label sparsity by leveraging label representations to enrich the labeled term set. Numerical experiments corroborate the effectiveness of our solution. In the third part, we introduce the concept of \\emph{Context-Aware Semantic Online Analytical Processing} (\\ie \\emph{CASeOLAP}) in text cubes, and use \\emph{top-$k$ representative phrases} to represent the semantics of the document subset in a text cube cell. By ranking phrases with a newly proposed ranking measure according to three criteria: integrity, popularity and distinctiveness. We identify phrases that can successfully digest the main content of a subset of documents of interest and contrast with other neighboring subsets. Our experiments in a large news dataset demonstrate the effectiveness of the newly proposed ranking measure in finding representative phrases and the efficiency in both query processing time and storage cost. The approach is also applied to clinical biomarker analysis and protest news analysis with success. In the last part, the system of \\emph{EventCube} is proposed to support end-to-end pipeline of text cube in an informative, interactive, and user-friendly manner. The system serves as a general platform for construction, search, summarization, OLAP (online analytical processing) and data mining on integrated text and structured data. The system is a growing testbed for various text cube based research and has been successfully applied to NASA for aviation safety report analysis and Army Research Lab for Counter-Terrorism Report analysis. To summarize, this thesis provides important results of construction and consumption of multi-dimensional text cubes and shows its power in tackling real-world text analysis tasks.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-12-01","The student, Fangbo Tao, accepted the attached license on 2017-12-06 at 11:45.","The student, Fangbo Tao, submitted this Dissertation for approval on 2017-12-06 at 11:53.","This Dissertation was approved for publication on 2017-12-06 at 13:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11883 on 2018-03-13 at 09:57:28","Made available in DSpace on 2018-03-13T15:25:31Z (GMT). No. of bitstreams: 2 TAO-DISSERTATION-2017.pdf: 7757301 bytes, checksum: eb265911f521a46b4473b8c6145b02e0 (MD5) LICENSE.txt: 4207 bytes, checksum: 441a8dcbebe919d9cdf10375f5c5b7ae (MD5) Previous issue date: 2017-12-06","Embargo set by: Seth Robbins for item 105207 Lift date: 2020-03-13T15:25:40Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 105207 Lift date: 2020-03-13T15:28:52Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 105207 on 2020-03-14T09:15:25Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/99244"],"dc:language":["en"],"dc:rights":["Copyright 2017 Fangbo Tao"],"dc:subject":["Text cube","Data cube","Data mining","Natural language processing","Text classification"],"dc:title":["Text cube: construction, summarization and mining"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:37Z"}