University of Cambridge
Adversarial Attacks on Natural Language and Speech Processing Models
Abstract
dc:description.abstractDeep learning models are parametric functions designed to capture patterns in data. The increasing availability of computational resources in recent years has facilitated the training and deployment of deep learning models across a wide range of applications in computer vision, natural language processing, and speech processing. Despite their high performance, deep learning models are vulnerable to adversarial attacks. A deliberate and specific perturbation of a clean input sample can create an adversarial example, which, when processed by the model, leads to incorrect predictions. Such vulnerabilities can be exploited by malicious adversaries in various high-stakes settings. For instance, a candidate could manipulate a written essay to deceive an automated scoring model into awarding an unmerited high grade. This thesis investigates the threat of adversarial attacks in deep learning systems, with a particular focus on natural language processing (NLP) and speech processing systems. The thesis begins by introducing fundamental deep learning concepts, including model architectures, training methodologies, and applications. This foundational discussion facilitates the understanding of modern systems, such as Transformer encoder-only language models (e.g., BERT) and generative large language models (LLMs) that use a Transformer decoder-only structure like ChatGPT and LLaMA. Subsequently, the theory underpinning adversarial attacks is presented, demonstrating that any attack can be characterised by three fundamental components: the attack goal, the constraints on adversarial perturbations, and the search process used to generate adversarial examples. These components are applied to describe specific attack methods across various input domains, ranging from the continuous space of images and speech to the discrete and sequential space of natural language tokens. The remainder of the thesis addresses challenges and solutions across diverse contexts. First, the attributes of adversarial attack methods in NLP are analysed, revealing their influence on defence strategies for classification tasks. Notably, it is shown that heavy miscalibration in encoder-only classifiers can obstruct the adversarial search process, leading to an overestimation of the effectiveness of adversarial training defence methods. Additionally, it is demonstrated that adversarial examples leave detectable residues in model encodings, which can be exploited for attack detection. Further discussions focus on defining adversarial attacks for generative NLP models. The complex nature of sequence outputs in generative tasks necessitates a sophisticated definition of adversarial attack goals to encompass diverse adversarial applications. A perception-based framework is proposed as a suitable approach to describe adversarial attacks across a variety of generative tasks. This framework is successfully applied to attacks on generative tasks such as Neural Machine Translation and Grammatical Error Correction. Further discussions focus on the use of an automated scoring module that operates on output sequences in the perception-based framework. It is shown that if powerful LLMs (such as ChatGPT) are leveraged for this scoring (LLM-as-a-Judge), an adversary can manipulate the sequence to artificially boost their score. The perception-based framework is extended to adversarial attacks in spoken language processing (SLP) systems. Traditional SLP tasks often employed a cascaded structure, with automatic speech recognition followed by downstream processing. However, recent flexible models, such as Whisper, enable end-to-end processing of multiple SLP tasks (e.g., speech transcription and translation) with the same model, using control prompts to guide the generation. Results highlight that adversaries can develop universal audio adversarial attacks exploiting non-acoustic tokens in the control prompt to manipulate SLP system behaviour. Finally, the thesis examines adversarial attacks from the perspective of individual data samples. A definition of sample attackability is proposed, identifying data input samples more susceptible to perturbations that produce adversarial examples. Empirical results demonstrate the effectiveness of a deep learning-based attackability detector in identifying highly attackable samples in both image and NLP contexts. In summary, the growing adoption of LLMs and speech-enabled LLMs for data processing and task automation heightens the importance of addressing adversarial threats. This thesis contributes towards enabling the safe deployment of these technologies in the future.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Raina, Vyas
- Advisor dc:contributor.advisor
-
- Gales, Mark
Subjects
dc:subject × 4Rights
dc:rightsIdentifiers
dc:identifier.*- Author Identifier
- 0000-0002-7157-5392
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/391167