Back to results

University of Illinois at Urbana-Champaign

TextGuard: Provable defense against backdoor attacks on text classification

Abstract

dc:description

Backdoor attacks against text classification have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against these backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong or adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. At a high level, TextGuard follows the partition and ensemble mechanism and first divides the (backdoored) training data into disjoint sub-training sets, where most subsets do not contain the backdoor trigger. Then, it trains multiple base classifiers from these subsets and ensemble them as the final classifier. TextGuard could guarantee the cleanliness of the majority base models and thus guarantee its prediction is unaffected by the backdoor trigger in both training and testing inputs. We conduct extensive evaluations for TextGuard on three benchmark datasets, provably and empirically, and demonstrate its superiority against state-of-the-art defenses under multiple backdoor attacks.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pei, Hengzhi
Contributors dc:contributor
  • Li, Bo

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Hengzhi Pei
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/120348

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Pei, Hengzhi. TextGuard: Provable defense against backdoor attacks on text classification. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/120348