University of Illinois at Urbana-Champaign
TextGuard: Provable defense against backdoor attacks on text classification
Abstract
dc:descriptionBackdoor attacks against text classification have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against these backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong or adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. At a high level, TextGuard follows the partition and ensemble mechanism and first divides the (backdoored) training data into disjoint sub-training sets, where most subsets do not contain the backdoor trigger. Then, it trains multiple base classifiers from these subsets and ensemble them as the final classifier. TextGuard could guarantee the cleanliness of the majority base models and thus guarantee its prediction is unaffected by the backdoor trigger in both training and testing inputs. We conduct extensive evaluations for TextGuard on three benchmark datasets, provably and empirically, and demonstrate its superiority against state-of-the-art defenses under multiple backdoor attacks.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Pei, Hengzhi
- Contributors dc:contributor
-
- Li, Bo
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Hengzhi Pei
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/120348