{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110579"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110579","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Mining social media stimulus from news article text using weakly-supervised narrative classification","abstract":"To make an accurate simulation for social media, we first need to find the stimulus in external sources. In this work, we model the stimulus mining into a narrative classification task on a news article dataset. The previous state-of-the-art text classification methods can not be directly applied here, mainly due to the following challenges we need to solve: 1) Lack of training data: the given news article data does not have labeling for narratives and we can not afford manual labeling other than a small evaluation set. 2) The complexity in narratives: narratives are defined in a more complex way comparing to the classes used in a classical news classification dataset, which stops us from using existing weakly supervised text classification methods that heavily depend on class name semantics. 3) The noisy news article dataset: the collected dataset does not guarantee the documents will belong to any of the narratives. In such cases, the power of the self-training strategy widely used in existing methods on weak supervision will be limited. To solve these challenges, we proposed a narrative decomposition and re-grouping strategy and a relevance filtering module, to fully utilize the power of weakly supervised classification methods. We conduct extensive experiments on two datasets under the background of real global events and further proposed two ways to combine different results for an optimal stimulus time-series.","abstract_html":"To make an accurate simulation for social media, we first need to find the stimulus in external sources. In this work, we model the stimulus mining into a narrative classification task on a news article dataset. The previous state-of-the-art text classification methods can not be directly applied here, mainly due to the following challenges we need to solve: 1) Lack of training data: the given news article data does not have labeling for narratives and we can not afford manual labeling other than a small evaluation set. 2) The complexity in narratives: narratives are defined in a more complex way comparing to the classes used in a classical news classification dataset, which stops us from using existing weakly supervised text classification methods that heavily depend on class name semantics. 3) The noisy news article dataset: the collected dataset does not guarantee the documents will belong to any of the narratives. In such cases, the power of the self-training strategy widely used in existing methods on weak supervision will be limited. To solve these challenges, we proposed a narrative decomposition and re-grouping strategy and a relevance filtering module, to fully utilize the power of weakly supervised classification methods. We conduct extensive experiments on two datasets under the background of real global events and further proposed two ways to combine different results for an optimal stimulus time-series.","abstract_has_math":false,"creators":["Qiu, Wenda"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T01:13:28Z","date_published":"2021-09-17T01:13:28Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Text Mining","News Classification"],"languages":["en"],"rights":["Copyright 2021 Wenda Qiu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110579","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Qiu, Wenda"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T01:13:28Z","2021-04-27","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Text Mining","News Classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Wenda Qiu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110579"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["To make an accurate simulation for social media, we first need to find the stimulus in external sources. In this work, we model the stimulus mining into a narrative classification task on a news article dataset. The previous state-of-the-art text classification methods can not be directly applied here, mainly due to the following challenges we need to solve: 1) Lack of training data: the given news article data does not have labeling for narratives and we can not afford manual labeling other than a small evaluation set. 2) The complexity in narratives: narratives are defined in a more complex way comparing to the classes used in a classical news classification dataset, which stops us from using existing weakly supervised text classification methods that heavily depend on class name semantics. 3) The noisy news article dataset: the collected dataset does not guarantee the documents will belong to any of the narratives. In such cases, the power of the self-training strategy widely used in existing methods on weak supervision will be limited. To solve these challenges, we proposed a narrative decomposition and re-grouping strategy and a relevance filtering module, to fully utilize the power of weakly supervised classification methods. We conduct extensive experiments on two datasets under the background of real global events and further proposed two ways to combine different results for an optimal stimulus time-series.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Wenda Qiu, accepted the attached license on 2021-04-26 at 15:48.","The student, Wenda Qiu, submitted this Thesis for approval on 2021-04-26 at 16:10.","This Thesis was approved for publication on 2021-04-27 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16566 on 2021-09-16 at 16:48:14","Made available in DSpace on 2021-09-17T01:13:28Z (GMT). No. of bitstreams: 3 QIU-THESIS-2021.pdf: 1129159 bytes, checksum: 94e33a669f0dcb2e3016e7e0fb756bb9 (MD5) Mining Social Media Stimulus from News Article Text using Weakly-supervised Narrative Classification (1).zip: 1206636 bytes, checksum: 8beb990a941fb5b831af59ca4d35eb07 (MD5) LICENSE.txt: 4206 bytes, checksum: 68940e80c66075d5a0100381d7eb13db (MD5) Previous issue date: 2021-04-27"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Mining social media stimulus from news article text using weakly-supervised narrative classification"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Qiu, Wenda"],"dc:date":["2021-09-17T01:13:28Z","2021-04-27","2021-05"],"dc:description":["To make an accurate simulation for social media, we first need to find the stimulus in external sources. In this work, we model the stimulus mining into a narrative classification task on a news article dataset. The previous state-of-the-art text classification methods can not be directly applied here, mainly due to the following challenges we need to solve: 1) Lack of training data: the given news article data does not have labeling for narratives and we can not afford manual labeling other than a small evaluation set. 2) The complexity in narratives: narratives are defined in a more complex way comparing to the classes used in a classical news classification dataset, which stops us from using existing weakly supervised text classification methods that heavily depend on class name semantics. 3) The noisy news article dataset: the collected dataset does not guarantee the documents will belong to any of the narratives. In such cases, the power of the self-training strategy widely used in existing methods on weak supervision will be limited. To solve these challenges, we proposed a narrative decomposition and re-grouping strategy and a relevance filtering module, to fully utilize the power of weakly supervised classification methods. We conduct extensive experiments on two datasets under the background of real global events and further proposed two ways to combine different results for an optimal stimulus time-series.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Wenda Qiu, accepted the attached license on 2021-04-26 at 15:48.","The student, Wenda Qiu, submitted this Thesis for approval on 2021-04-26 at 16:10.","This Thesis was approved for publication on 2021-04-27 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16566 on 2021-09-16 at 16:48:14","Made available in DSpace on 2021-09-17T01:13:28Z (GMT). No. of bitstreams: 3 QIU-THESIS-2021.pdf: 1129159 bytes, checksum: 94e33a669f0dcb2e3016e7e0fb756bb9 (MD5) Mining Social Media Stimulus from News Article Text using Weakly-supervised Narrative Classification (1).zip: 1206636 bytes, checksum: 8beb990a941fb5b831af59ca4d35eb07 (MD5) LICENSE.txt: 4206 bytes, checksum: 68940e80c66075d5a0100381d7eb13db (MD5) Previous issue date: 2021-04-27"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110579"],"dc:language":["en"],"dc:rights":["Copyright 2021 Wenda Qiu"],"dc:subject":["Text Mining","News Classification"],"dc:title":["Mining social media stimulus from news article text using weakly-supervised narrative classification"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}