{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125539"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125539","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Weakly supervised text mining with text-rich networks","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_has_math":false,"creators":["Zhang, Xinyang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Sundaram, Hari","Tong, Hanghang","Dong, Xin Luna"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-06-27","date_published":"2024-06-27","updated_at":"2026-07-22T22:25:02Z","subjects":["Language Models","Text Mining","Text-rich Networks","Weak Supervision"],"languages":["en","eng"],"rights":["Copyright 2024, Xinyang Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125539","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Sundaram, Hari","Tong, Hanghang","Dong, Xin Luna"]},{"key":"dc:creator","label":"Author","values":["Zhang, Xinyang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-06-27","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Language Models","Text Mining","Text-rich Networks","Weak Supervision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024, Xinyang Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125539"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Xinyang Zhang, accepted the attached license on 2024-06-25 at 23:58.","The student, Xinyang Zhang, submitted this Dissertation for approval on 2024-06-26 at 00:05.","This Dissertation was approved for publication on 2024-06-27 at 12:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20876 on 2025-02-04 at 21:03:56","The advent of the information age has brought about an unprecedented surge in the volume of data available at our fingertips. A significant portion of this data is manifested in the form of semi-structured text corpora, which are characterized by the interplay between unstructured text and structured metadata. These corpora are ubiquitous across various domains, presenting unique challenges for knowledge discovery and data mining. For instance, social media platforms are a complex blend of users, user-generated content, and the intricate web of interactions between them. Similarly, academic publication databases comprise a rich tapestry of published content, author information, and venue details. The sheer scale and heterogeneity of these semi-structured text corpora necessitate the development of intelligent and efficient techniques to extract, organize, and summarize the valuable knowledge embedded within them. Traditional approaches to text mining often rely heavily on extensive human annotation or manually curated knowledge bases, which are not only time-consuming and expensive but also difficult to adapt to new domains. To fully harness the potential of these vast information resources, we propose a novel approach that leverages the inherent structure of the data itself. We introduce a framework based on text-rich networks, which effectively integrates unstructured documents with structured metadata to create a comprehensive representation for data mining. By interconnecting documents based on shared attributes, such as linking research papers authored by the same individual, we construct a unified network that captures both the textual content and the relationships between entities. This approach opens up new possibilities for knowledge discovery, enabling us to uncover valuable patterns and insights that may be difficult to detect using traditional methods. The text-rich network provides a natural encoding of document relevance, allowing us to develop weakly supervised text mining applications that require minimal human intervention. This methodology reduces the reliance on large amounts of labeled data, making it more scalable and adaptable to various domains. Through the use of text-rich networks, we aim to enhance the efficiency and effectiveness of knowledge discovery in semi-structured text corpora, facilitating data-driven insights and innovation in the field of text mining. My research consists of two main areas of investigation: (1) construction and consolidation of a text-rich network, and (2) mining of a text-rich network. In the first area, we focus on building a robust text-rich network structure that relies on the availability of structured information alongside text. When such information is unavailable or incomplete, we investigate weakly supervised methods to extract structured information. 1. Open-World Attribute Value Extraction: We propose a novel approach for open-world attribute value extraction, which enables us to identify entities and attributes that facilitate the construction of a comprehensive text-rich network. This method addresses the challenges of incomplete or missing structured information in semi-structured text corpora. In the second area, we explore various methods built on top of a fully constructed text-rich network for fundamental text representation and weakly supervised text mining applications. 2. Language Model Pre-training with Text-Rich Networks: We introduce a novel language model pre-training approach that utilizes the text-rich network structure to capture more meaningful representations of text data. By incorporating the network’s contextual information during the pre-training process, we aim to improve the effectiveness of downstream natural language processing tasks. 3. Text-Rich Network-Driven Weakly Supervised Text Classification: We develop a framework that leverages the rich structural information embedded in the text-rich network to enhance the performance of text classification tasks. This approach reduces the reliance on large amounts of labeled data, making it more adaptable to real-world scenarios. Together, these components create a cohesive framework for weakly supervised text mining with text-rich networks. Our open-world attribute value extraction method strengthens the construction and consolidation of text-rich networks, while our text classification and language model pre-training approaches demonstrate the power of mining these networks for various applications. By integrating these techniques, we establish a comprehensive methodology that enables efficient and effective knowledge discovery in semi-structured text corpora, reducing the reliance on human annotation and making it more adaptable to real-world scenarios."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Weakly supervised text mining with text-rich networks"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Sundaram, Hari","Tong, Hanghang","Dong, Xin Luna"],"dc:creator":["Zhang, Xinyang"],"dc:date":["2024-06-27","2024-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Xinyang Zhang, accepted the attached license on 2024-06-25 at 23:58.","The student, Xinyang Zhang, submitted this Dissertation for approval on 2024-06-26 at 00:05.","This Dissertation was approved for publication on 2024-06-27 at 12:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20876 on 2025-02-04 at 21:03:56","The advent of the information age has brought about an unprecedented surge in the volume of data available at our fingertips. A significant portion of this data is manifested in the form of semi-structured text corpora, which are characterized by the interplay between unstructured text and structured metadata. These corpora are ubiquitous across various domains, presenting unique challenges for knowledge discovery and data mining. For instance, social media platforms are a complex blend of users, user-generated content, and the intricate web of interactions between them. Similarly, academic publication databases comprise a rich tapestry of published content, author information, and venue details. The sheer scale and heterogeneity of these semi-structured text corpora necessitate the development of intelligent and efficient techniques to extract, organize, and summarize the valuable knowledge embedded within them. Traditional approaches to text mining often rely heavily on extensive human annotation or manually curated knowledge bases, which are not only time-consuming and expensive but also difficult to adapt to new domains. To fully harness the potential of these vast information resources, we propose a novel approach that leverages the inherent structure of the data itself. We introduce a framework based on text-rich networks, which effectively integrates unstructured documents with structured metadata to create a comprehensive representation for data mining. By interconnecting documents based on shared attributes, such as linking research papers authored by the same individual, we construct a unified network that captures both the textual content and the relationships between entities. This approach opens up new possibilities for knowledge discovery, enabling us to uncover valuable patterns and insights that may be difficult to detect using traditional methods. The text-rich network provides a natural encoding of document relevance, allowing us to develop weakly supervised text mining applications that require minimal human intervention. This methodology reduces the reliance on large amounts of labeled data, making it more scalable and adaptable to various domains. Through the use of text-rich networks, we aim to enhance the efficiency and effectiveness of knowledge discovery in semi-structured text corpora, facilitating data-driven insights and innovation in the field of text mining. My research consists of two main areas of investigation: (1) construction and consolidation of a text-rich network, and (2) mining of a text-rich network. In the first area, we focus on building a robust text-rich network structure that relies on the availability of structured information alongside text. When such information is unavailable or incomplete, we investigate weakly supervised methods to extract structured information. 1. Open-World Attribute Value Extraction: We propose a novel approach for open-world attribute value extraction, which enables us to identify entities and attributes that facilitate the construction of a comprehensive text-rich network. This method addresses the challenges of incomplete or missing structured information in semi-structured text corpora. In the second area, we explore various methods built on top of a fully constructed text-rich network for fundamental text representation and weakly supervised text mining applications. 2. Language Model Pre-training with Text-Rich Networks: We introduce a novel language model pre-training approach that utilizes the text-rich network structure to capture more meaningful representations of text data. By incorporating the network’s contextual information during the pre-training process, we aim to improve the effectiveness of downstream natural language processing tasks. 3. Text-Rich Network-Driven Weakly Supervised Text Classification: We develop a framework that leverages the rich structural information embedded in the text-rich network to enhance the performance of text classification tasks. This approach reduces the reliance on large amounts of labeled data, making it more adaptable to real-world scenarios. Together, these components create a cohesive framework for weakly supervised text mining with text-rich networks. Our open-world attribute value extraction method strengthens the construction and consolidation of text-rich networks, while our text classification and language model pre-training approaches demonstrate the power of mining these networks for various applications. By integrating these techniques, we establish a comprehensive methodology that enables efficient and effective knowledge discovery in semi-structured text corpora, reducing the reliance on human annotation and making it more adaptable to real-world scenarios."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125539"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024, Xinyang Zhang"],"dc:subject":["Language Models","Text Mining","Text-rich Networks","Weak Supervision"],"dc:title":["Weakly supervised text mining with text-rich networks"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}