{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104867"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104867","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Weakly-supervised text classification","abstract":"Deep neural networks are gaining increasing popularity for the classic text classification task, due to their strong expressive power and less requirement for feature engineering. Despite such attractiveness, neural text classification models suffer from the lack of training data in many real-world applications. Although many semi-supervised and weakly-supervised text classification models exist, they cannot be easily applied to deep neural models and meanwhile support limited supervision types. In this work, we propose a weakly-supervised framework that addresses the lack of training data in neural text classification. Our framework consists of two modules: (1) a pseudo-document generator that leverages seed information to generate pseudo-labeled documents for model pre-training, and (2) a self-training module that bootstraps on real unlabeled data for model refinement. Our framework has the flexibility to handle different types of weak supervision and can be easily integrated into existing deep neural models for text classification. Based on this framework, we propose two methods, WeSTClass and WeSHClass, for flat text classification and hierarchical text classification, respectively. We have performed extensive experiments on real-world datasets from different domains. The results demonstrate that our proposed framework achieves inspiring performance without requiring excessive training data and outperforms baselines significantly.","abstract_html":"Deep neural networks are gaining increasing popularity for the classic text classification task, due to their strong expressive power and less requirement for feature engineering. Despite such attractiveness, neural text classification models suffer from the lack of training data in many real-world applications. Although many semi-supervised and weakly-supervised text classification models exist, they cannot be easily applied to deep neural models and meanwhile support limited supervision types. In this work, we propose a weakly-supervised framework that addresses the lack of training data in neural text classification. Our framework consists of two modules: (1) a pseudo-document generator that leverages seed information to generate pseudo-labeled documents for model pre-training, and (2) a self-training module that bootstraps on real unlabeled data for model refinement. Our framework has the flexibility to handle different types of weak supervision and can be easily integrated into existing deep neural models for text classification. Based on this framework, we propose two methods, WeSTClass and WeSHClass, for flat text classification and hierarchical text classification, respectively. We have performed extensive experiments on real-world datasets from different domains. The results demonstrate that our proposed framework achieves inspiring performance without requiring excessive training data and outperforms baselines significantly.","abstract_has_math":false,"creators":["Meng, Yu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T19:55:46Z","date_published":"2019-08-23T19:55:46Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Text Classification","Weakly-supervised Learning","Neural Classification Model","Hierarchical Classification"],"languages":["en"],"rights":["Copyright 2019 Yu Meng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104867","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Meng, Yu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T19:55:46Z","2019-04-22","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Text Classification","Weakly-supervised Learning","Neural Classification Model","Hierarchical Classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yu Meng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104867"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Deep neural networks are gaining increasing popularity for the classic text classification task, due to their strong expressive power and less requirement for feature engineering. Despite such attractiveness, neural text classification models suffer from the lack of training data in many real-world applications. Although many semi-supervised and weakly-supervised text classification models exist, they cannot be easily applied to deep neural models and meanwhile support limited supervision types. In this work, we propose a weakly-supervised framework that addresses the lack of training data in neural text classification. Our framework consists of two modules: (1) a pseudo-document generator that leverages seed information to generate pseudo-labeled documents for model pre-training, and (2) a self-training module that bootstraps on real unlabeled data for model refinement. Our framework has the flexibility to handle different types of weak supervision and can be easily integrated into existing deep neural models for text classification. Based on this framework, we propose two methods, WeSTClass and WeSHClass, for flat text classification and hierarchical text classification, respectively. We have performed extensive experiments on real-world datasets from different domains. The results demonstrate that our proposed framework achieves inspiring performance without requiring excessive training data and outperforms baselines significantly.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yu Meng, accepted the attached license on 2019-04-18 at 12:45.","The student, Yu Meng, submitted this Thesis for approval on 2019-04-18 at 12:52.","This Thesis was approved for publication on 2019-04-22 at 15:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13745 on 2019-08-22 at 14:44:48","Made available in DSpace on 2019-08-23T19:55:46Z (GMT). No. of bitstreams: 2 MENG-THESIS-2019.pdf: 1354812 bytes, checksum: 1829017018ca6d79ab8c4d82d1ebd5f4 (MD5) LICENSE.txt: 4204 bytes, checksum: 4ea4555dd29f4c31eddc669378b1c9d3 (MD5) Previous issue date: 2019-04-22"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Weakly-supervised text classification"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Meng, Yu"],"dc:date":["2019-08-23T19:55:46Z","2019-04-22","2019-05"],"dc:description":["Deep neural networks are gaining increasing popularity for the classic text classification task, due to their strong expressive power and less requirement for feature engineering. Despite such attractiveness, neural text classification models suffer from the lack of training data in many real-world applications. Although many semi-supervised and weakly-supervised text classification models exist, they cannot be easily applied to deep neural models and meanwhile support limited supervision types. In this work, we propose a weakly-supervised framework that addresses the lack of training data in neural text classification. Our framework consists of two modules: (1) a pseudo-document generator that leverages seed information to generate pseudo-labeled documents for model pre-training, and (2) a self-training module that bootstraps on real unlabeled data for model refinement. Our framework has the flexibility to handle different types of weak supervision and can be easily integrated into existing deep neural models for text classification. Based on this framework, we propose two methods, WeSTClass and WeSHClass, for flat text classification and hierarchical text classification, respectively. We have performed extensive experiments on real-world datasets from different domains. The results demonstrate that our proposed framework achieves inspiring performance without requiring excessive training data and outperforms baselines significantly.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yu Meng, accepted the attached license on 2019-04-18 at 12:45.","The student, Yu Meng, submitted this Thesis for approval on 2019-04-18 at 12:52.","This Thesis was approved for publication on 2019-04-22 at 15:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13745 on 2019-08-22 at 14:44:48","Made available in DSpace on 2019-08-23T19:55:46Z (GMT). No. of bitstreams: 2 MENG-THESIS-2019.pdf: 1354812 bytes, checksum: 1829017018ca6d79ab8c4d82d1ebd5f4 (MD5) LICENSE.txt: 4204 bytes, checksum: 4ea4555dd29f4c31eddc669378b1c9d3 (MD5) Previous issue date: 2019-04-22"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104867"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yu Meng"],"dc:subject":["Text Classification","Weakly-supervised Learning","Neural Classification Model","Hierarchical Classification"],"dc:title":["Weakly-supervised text classification"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}