{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120348"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120348","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"TextGuard: Provable defense against backdoor attacks on text classification","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Pei, Hengzhi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Li, Bo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Backdoor Attack","Text Classification","Provable Defense"],"languages":["en","eng"],"rights":["Copyright 2023 Hengzhi Pei"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120348","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Li, Bo"]},{"key":"dc:creator","label":"Author","values":["Pei, Hengzhi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-04-06"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Backdoor Attack","Text Classification","Provable Defense"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Hengzhi Pei"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120348"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Hengzhi Pei, accepted the attached license on 2023-04-03 at 17:17.","The student, Hengzhi Pei, submitted this Thesis for approval on 2023-04-03 at 17:31.","This Thesis was approved for publication on 2023-04-06 at 14:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18916 on 2023-09-01 at 17:13:08","Backdoor attacks against text classification have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against these backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong or adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. At a high level, TextGuard follows the partition and ensemble mechanism and first divides the (backdoored) training data into disjoint sub-training sets, where most subsets do not contain the backdoor trigger. Then, it trains multiple base classifiers from these subsets and ensemble them as the final classifier. TextGuard could guarantee the cleanliness of the majority base models and thus guarantee its prediction is unaffected by the backdoor trigger in both training and testing inputs. We conduct extensive evaluations for TextGuard on three benchmark datasets, provably and empirically, and demonstrate its superiority against state-of-the-art defenses under multiple backdoor attacks."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["TextGuard: Provable defense against backdoor attacks on text classification"]}]}],"canonical_facts":{"dc:contributor":["Li, Bo"],"dc:creator":["Pei, Hengzhi"],"dc:date":["2023-05","2023-04-06"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Hengzhi Pei, accepted the attached license on 2023-04-03 at 17:17.","The student, Hengzhi Pei, submitted this Thesis for approval on 2023-04-03 at 17:31.","This Thesis was approved for publication on 2023-04-06 at 14:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18916 on 2023-09-01 at 17:13:08","Backdoor attacks against text classification have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against these backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong or adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. At a high level, TextGuard follows the partition and ensemble mechanism and first divides the (backdoored) training data into disjoint sub-training sets, where most subsets do not contain the backdoor trigger. Then, it trains multiple base classifiers from these subsets and ensemble them as the final classifier. TextGuard could guarantee the cleanliness of the majority base models and thus guarantee its prediction is unaffected by the backdoor trigger in both training and testing inputs. We conduct extensive evaluations for TextGuard on three benchmark datasets, provably and empirically, and demonstrate its superiority against state-of-the-art defenses under multiple backdoor attacks."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120348"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Hengzhi Pei"],"dc:subject":["Backdoor Attack","Text Classification","Provable Defense"],"dc:title":["TextGuard: Provable defense against backdoor attacks on text classification"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}