{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115488"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115488","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Annotation-free knowledge mining from massive text corpora","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Gu, Xiaotao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yu, Cong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["text mining","knowledge mining"],"languages":["en","eng"],"rights":["Copyright 2022 Xiaotao Gu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115488","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yu, Cong"]},{"key":"dc:creator","label":"Author","values":["Gu, Xiaotao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-19"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["text mining","knowledge mining"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Xiaotao Gu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115488"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Xiaotao Gu, accepted the attached license on 2022-04-19 at 00:14.","The student, Xiaotao Gu, submitted this Dissertation for approval on 2022-04-19 at 00:27.","This Dissertation was approved for publication on 2022-04-19 at 13:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17778 on 2022-11-11 at 13:15:42","Recent years have witnessed unprecedented information explosion. Text data, as a major carrier of information and knowledge, is emerging in blazing speed thanks to the development of the Internet. Despite its great value, the overwhelming amount of text data brings great challenges, both for human readers to consume and for machines to process. To fully unleash the power of text data, our goal is to mine, organize and summarize structured knowledge from massive text corpora in an intelligent and effort-light manner. Existing models for text mining heavily lean on excessive human annotation, or manually curated knowledge bases, which are not only expensive, but also hard to transfer to new domains. In this work, we aim to alleviate the need for human annotation, where we directly mine high-quality supervision signals from the input corpora for model learning. We show that through self-supervised tasks on the corpora, models can effectively capture rich patterns in text, and learn to extract information and generate text with the guidance of such patterns as silver labels. In this thesis, we outline a general framework to mine the knowledge pyramid of knowledge from different granularities to satisfy various information needs. We explore the possibility of annotation-free knowledge mining on phrase-level, sentence-level, document-level, and multi-document-level tasks: 1. Automated Phrase Mining. Phrases are arguably the basic semantic unit for text understanding. We start from introducing UCPhrase, an unsupervised context-aware phrase tagging model, which does not rely on any human annotation or knowledge bases. We show that silver labels mined from unlabeled corpora can replace and even outperform distantly fetched labels from existing knowledge bases. We further propose to leverage attention maps generated by pre-trained language models to extract informative features about sentence structures, which alleviate the frequency bias and can effectively capture emerging infrequent phrases. 2. Phrase-aware Sentence Parsing. With the knowledge about phrases, we further study how phrases are connected to form the structure of a sentence. We propose to model sentence parsing as a three-stage process: (1) extract obvious phrases (e.g., entities, names, concepts) in the sentence as prior knowledge; (2) learn to identify more local phrases with the guidance from extracted phrases; (3) learn to connect local phrases to form high-level structures. We treat randomly masked tokens in sentences as silver labels, and incorporate phrase information into structured language models (LM) for unsupervised constituency parsing. The proposed phrase-regularized warm-up and phrase-aware masked language modeling improve both local and high-level structure parsing, and establish the new SOTA for LM-based constituency parsing. 3. Representative Headline Generation. Beyond syntactic knowledge extraction, we generalize the idea of annotation-free mining to automatically organize and summarize documents based on semantic knowledge. With the real need from news readers to efficiently consume overwhelming daily news, we develop NHNet, a self-supervised model to generate concise and high-quality headlines for news stories. Without human annotation, we propose a three-level pre-training framework to fully leverage silver labels from web-scale news data for model training. We show that our model outperforms supervised generation model trained on human labels collected for years, and shows even stronger performance with finetuning on a small amount of human labels. 4. Information-guided Document Summarization. We then demonstrate how to incorporate phrase and sentence-level knowledge into self-supervised generation models for more accurate and interpretable document summarization. We propose EASum, a two-stage framework for unsupervised abstractive summarization. We train an information-guided generation model to generate fluent sentences covering specified phrases and events with silver labels mined from unlabeled documents. The model then generates a summary for each document with key phrases and events extracted from the target document. We show that key phrases and parsing-based events improve the accuracy of generated summaries, and also leave a clear clue for understanding the source of the generated summary. Together, the developed methods form a powerful annotation-free framework for multi-level text mining. All models mentioned above are open sourced for public usage, and some have been deployed in real-world production systems."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Annotation-free knowledge mining from massive text corpora"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yu, Cong"],"dc:creator":["Gu, Xiaotao"],"dc:date":["2022-05","2022-04-19"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Xiaotao Gu, accepted the attached license on 2022-04-19 at 00:14.","The student, Xiaotao Gu, submitted this Dissertation for approval on 2022-04-19 at 00:27.","This Dissertation was approved for publication on 2022-04-19 at 13:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17778 on 2022-11-11 at 13:15:42","Recent years have witnessed unprecedented information explosion. Text data, as a major carrier of information and knowledge, is emerging in blazing speed thanks to the development of the Internet. Despite its great value, the overwhelming amount of text data brings great challenges, both for human readers to consume and for machines to process. To fully unleash the power of text data, our goal is to mine, organize and summarize structured knowledge from massive text corpora in an intelligent and effort-light manner. Existing models for text mining heavily lean on excessive human annotation, or manually curated knowledge bases, which are not only expensive, but also hard to transfer to new domains. In this work, we aim to alleviate the need for human annotation, where we directly mine high-quality supervision signals from the input corpora for model learning. We show that through self-supervised tasks on the corpora, models can effectively capture rich patterns in text, and learn to extract information and generate text with the guidance of such patterns as silver labels. In this thesis, we outline a general framework to mine the knowledge pyramid of knowledge from different granularities to satisfy various information needs. We explore the possibility of annotation-free knowledge mining on phrase-level, sentence-level, document-level, and multi-document-level tasks: 1. Automated Phrase Mining. Phrases are arguably the basic semantic unit for text understanding. We start from introducing UCPhrase, an unsupervised context-aware phrase tagging model, which does not rely on any human annotation or knowledge bases. We show that silver labels mined from unlabeled corpora can replace and even outperform distantly fetched labels from existing knowledge bases. We further propose to leverage attention maps generated by pre-trained language models to extract informative features about sentence structures, which alleviate the frequency bias and can effectively capture emerging infrequent phrases. 2. Phrase-aware Sentence Parsing. With the knowledge about phrases, we further study how phrases are connected to form the structure of a sentence. We propose to model sentence parsing as a three-stage process: (1) extract obvious phrases (e.g., entities, names, concepts) in the sentence as prior knowledge; (2) learn to identify more local phrases with the guidance from extracted phrases; (3) learn to connect local phrases to form high-level structures. We treat randomly masked tokens in sentences as silver labels, and incorporate phrase information into structured language models (LM) for unsupervised constituency parsing. The proposed phrase-regularized warm-up and phrase-aware masked language modeling improve both local and high-level structure parsing, and establish the new SOTA for LM-based constituency parsing. 3. Representative Headline Generation. Beyond syntactic knowledge extraction, we generalize the idea of annotation-free mining to automatically organize and summarize documents based on semantic knowledge. With the real need from news readers to efficiently consume overwhelming daily news, we develop NHNet, a self-supervised model to generate concise and high-quality headlines for news stories. Without human annotation, we propose a three-level pre-training framework to fully leverage silver labels from web-scale news data for model training. We show that our model outperforms supervised generation model trained on human labels collected for years, and shows even stronger performance with finetuning on a small amount of human labels. 4. Information-guided Document Summarization. We then demonstrate how to incorporate phrase and sentence-level knowledge into self-supervised generation models for more accurate and interpretable document summarization. We propose EASum, a two-stage framework for unsupervised abstractive summarization. We train an information-guided generation model to generate fluent sentences covering specified phrases and events with silver labels mined from unlabeled documents. The model then generates a summary for each document with key phrases and events extracted from the target document. We show that key phrases and parsing-based events improve the accuracy of generated summaries, and also leave a clear clue for understanding the source of the generated summary. Together, the developed methods form a powerful annotation-free framework for multi-level text mining. All models mentioned above are open sourced for public usage, and some have been deployed in real-world production systems."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115488"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Xiaotao Gu"],"dc:subject":["text mining","knowledge mining"],"dc:title":["Annotation-free knowledge mining from massive text corpora"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}