{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125546"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125546","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Guarding truthfulness: Detecting and correcting false information","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_has_math":false,"creators":["Huang, Kung-Hsiang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Zhai, ChengXiang","McKeown, Kathleen","Peng, Hao","Zhao, Han","Joty, Shafiq"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-08","date_published":"2024-08","updated_at":"2026-07-22T22:25:02Z","subjects":["Fact-checking","Factual Inconsistency","Fake News Detection","Factual Error Correction"],"languages":["en","eng"],"rights":["Copyright 2024 Kung-Hsiang Huang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125546","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Zhai, ChengXiang","McKeown, Kathleen","Peng, Hao","Zhao, Han","Joty, Shafiq"]},{"key":"dc:creator","label":"Author","values":["Huang, Kung-Hsiang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-08","2024-06-28"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Fact-checking","Factual Inconsistency","Fake News Detection","Factual Error Correction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Kung-Hsiang Huang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125546"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Kung-Hsiang Huang, accepted the attached license on 2024-06-27 at 16:36.","The student, Kung-Hsiang Huang, submitted this Dissertation for approval on 2024-06-27 at 16:45.","This Dissertation was approved for publication on 2024-06-28 at 14:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20889 on 2025-02-04 at 21:03:58","The proliferation of false information in today’s digital age is becoming a global epidemic, leading to societal instability and discord. Large language models, such as ChatGPT, have unintentionally escalated the problem by occasionally generating misleading content, promoting harmful views, and eroding faith in public conversation. Based on these realities, this thesis sets out to establish effective strategies to counter and curtail the spread of false information. We approach the mitigation of false news from two distinct angles: detection and correction. Robustly detecting false information is a stepping stone towards combating false information and preventing the further spread of inaccurate content. We have contributed to the field by proposing two approaches that address two important challenges: the gap between machine-generated fake news and human-written fake news, and the scarcity of fact-checking annotated data in low-resource languages. First, we generate more realistic fake news by explicitly incorporating propaganda techniques and adopting self-critical sequence training to instruct the generator not to produce factual content. We show that detectors trained on our generated data are more effective in identifying human-written fake news. Second, for the challenge of data scarcity, we leverage cross-lingual retrieval techniques to identify relevant passages across languages by learning to retrieve from different languages via pseudo-feedback, thereby enabling cross-lingual fact-checking. Our method outperforms existing models and alleviates the data scarcity issues. Upon identifying false information, the intuitive next step is to automatically correct the information based on the supporting or refuting evidence retrieved. To this end, we devised an interpretable factual error correction framework, which remediates factual errors in a zero-shot manner. This framework, inspired by how humans correct errors, asks questions about the input claim, looks for answers in the evidence, and reconstructs a factual statement. By breaking this task into several sub-tasks that have much richer resources, we eliminate the need for training data in factual error correction. More importantly, the decomposability of our framework offers natural interpretability, as the questions and answers generated explicitly illuminate which parts of the original claim were incorrect and why. We demonstrate that our approach outperforms fully-supervised approaches that have been fine-tuned. Additionally, we extended the scope of factual error correction to visual data, focusing on chart understanding. Recognizing the inherent complexity of interpreting graphical information, we proposed a methodology for converting charts into tables for enhanced error correction. Furthermore, we introduced the Multi-Agent Debate Refinement (MADR) framework to tackle unfaithfulness in explanations generated by large language models for fact-checking. This iterative, multi-agent strategy significantly enhanced the faithfulness and reliability of these explanations."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Guarding truthfulness: Detecting and correcting false information"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Zhai, ChengXiang","McKeown, Kathleen","Peng, Hao","Zhao, Han","Joty, Shafiq"],"dc:creator":["Huang, Kung-Hsiang"],"dc:date":["2024-08","2024-06-28"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Kung-Hsiang Huang, accepted the attached license on 2024-06-27 at 16:36.","The student, Kung-Hsiang Huang, submitted this Dissertation for approval on 2024-06-27 at 16:45.","This Dissertation was approved for publication on 2024-06-28 at 14:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20889 on 2025-02-04 at 21:03:58","The proliferation of false information in today’s digital age is becoming a global epidemic, leading to societal instability and discord. Large language models, such as ChatGPT, have unintentionally escalated the problem by occasionally generating misleading content, promoting harmful views, and eroding faith in public conversation. Based on these realities, this thesis sets out to establish effective strategies to counter and curtail the spread of false information. We approach the mitigation of false news from two distinct angles: detection and correction. Robustly detecting false information is a stepping stone towards combating false information and preventing the further spread of inaccurate content. We have contributed to the field by proposing two approaches that address two important challenges: the gap between machine-generated fake news and human-written fake news, and the scarcity of fact-checking annotated data in low-resource languages. First, we generate more realistic fake news by explicitly incorporating propaganda techniques and adopting self-critical sequence training to instruct the generator not to produce factual content. We show that detectors trained on our generated data are more effective in identifying human-written fake news. Second, for the challenge of data scarcity, we leverage cross-lingual retrieval techniques to identify relevant passages across languages by learning to retrieve from different languages via pseudo-feedback, thereby enabling cross-lingual fact-checking. Our method outperforms existing models and alleviates the data scarcity issues. Upon identifying false information, the intuitive next step is to automatically correct the information based on the supporting or refuting evidence retrieved. To this end, we devised an interpretable factual error correction framework, which remediates factual errors in a zero-shot manner. This framework, inspired by how humans correct errors, asks questions about the input claim, looks for answers in the evidence, and reconstructs a factual statement. By breaking this task into several sub-tasks that have much richer resources, we eliminate the need for training data in factual error correction. More importantly, the decomposability of our framework offers natural interpretability, as the questions and answers generated explicitly illuminate which parts of the original claim were incorrect and why. We demonstrate that our approach outperforms fully-supervised approaches that have been fine-tuned. Additionally, we extended the scope of factual error correction to visual data, focusing on chart understanding. Recognizing the inherent complexity of interpreting graphical information, we proposed a methodology for converting charts into tables for enhanced error correction. Furthermore, we introduced the Multi-Agent Debate Refinement (MADR) framework to tackle unfaithfulness in explanations generated by large language models for fact-checking. This iterative, multi-agent strategy significantly enhanced the faithfulness and reliability of these explanations."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125546"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Kung-Hsiang Huang"],"dc:subject":["Fact-checking","Factual Inconsistency","Fake News Detection","Factual Error Correction"],"dc:title":["Guarding truthfulness: Detecting and correcting false information"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}