University of Technology Sydney
Adversarial training objectives for generative attacks on text classifiers
Abstract
dc:description.abstractIn natural language processing, creating textual adversarial examples is challenging. These examples aim to deceive text classifiers into incorrect predictions while maintaining linguistic similarity to genuine inputs. This complexity arises from the need to preserve the original text's fluency, grammaticality, and semantic coherence, which complicates the generation of effective adversarial samples. This thesis has focused on refining generative models to produce textual adversarial examples, providing an alternative to the prevalent token-modification attacks that transform document tokens using methods such as word insertion and replacement. Generative approaches, which start with a pre-trained sequence-to-sequence (seq2seq) transformer model and fine-tune it to generate adversarial content, offer benefits like efficiency and parallel processing capabilities. Despite these advantages, generative methods remain under-explored and lack established training methodologies. This research addresses this gap by adapting generative models across various scenarios characterised by unique threat models, assumptions, and objectives, aiming to create examples that mislead classifiers while preserving essential textual properties. In our first black-box scenario, where only the victim model's predictions are known, our approach uses reinforcement learning with a policy gradient algorithm to fine-tune the seq2seq model, using a reward function that encourages adversarial examples and penalising constraint violations. In the next white-box scenario, the attacker has full knowledge of the victim model's architecture and parameters, accessing loss function gradients. Our method proposes an original differentiable objective function with pre-trained models to improve the textual quality of the adversarial examples, while inducing the victim model into an incorrect classification. Token-mapping matrices connect each model, ensuring end-to-end differentiability and allowing integration across vocabularies, overcoming a key limitation of prior work. The final scenario adapts our previous white-box attack to a multilingual context, requiring the generation of adversarial text in the same language as the original. Our approach, successfully demonstrated across five languages, integrates a language-detection model into the training objective and to our knowledge represents the first generative multilingual attack. In summary, this thesis has demonstrated the effectiveness of our generative methodologies, highlighting their potential to advance textual adversarial example creation and positioning them as a competitive alternative to established token-modification techniques.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Roth, Thomas Paul
Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- The author owns the copyright in this thesis including all reproduction and reuse rights for the work. The work may not be altered without the permission of the copyright owner. Attribution is essential when quoting or paraphrasing from this thesis.
- © 2024 Thomas Paul Roth
- au.edu.uts.lib/cph
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/10453/187457
- OAI identifier oai:identifier
- oai:opus.lib.uts.edu.au:10453/187457