Back to results

University of Illinois at Urbana-Champaign

Provable stability defenses for targeted data poisoning

Abstract

dc:description

Modern machine learning systems are often trained on massive, crowdsourced datasets. Due to the impossibility of checking this data, these systems may be susceptible to data poisoning attacks where malicious users inject false training data in order to influence the learned model. While recent work has focused primarily on the untargeted case, where the attacker's goal is to increase overall error, much less is understood about the theoretical underpinnings of targeted data poisoning attacks. These attacks try to cause the learned model to change its prediction on only a few targeted examples without raising suspicion. We suggest algorithmic stability as a sufficient condition for robustness against data poisoning, construct upper bounds on the possible effectiveness of data poisoning attacks against stable algorithms, and propose an algorithm that provides resilience against popular classes of attacks. Empirically, we report findings on the MNIST 1-7 image classification dataset and the TREC 2007 spam detection dataset that confirms our theoretical findings.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Vijitbenjaronk, Warut D.
Contributors dc:contributor
  • Koyejo, Oluwasanmi

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Copyright 2020 Warut D. Vijitbenjaronk
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/108638
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/108638

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Vijitbenjaronk, Warut D.. Provable stability defenses for targeted data poisoning. Thesis thesis, University of Illinois at Urbana-Champaign, 2020. http://hdl.handle.net/2142/108638