Virginia Tech
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
Abstract
dc:description.abstractLarge language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models.
Degree
thesis:*- Name thesis:degree_name
- Master of Science
- Level thesis:degree_level
- masters
- Discipline thesis:degree_discipline
- Computer Engineering
- Department dc:contributor.department
- Electrical and Computer Engineering
- Grantor dc:publisher
- Virginia Tech
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Dabas, Mahavir
- Chair dc:contributor.committeechair
-
- Jia, Ruoxi
- Committee members dc:contributor.committeemember
-
- Ramakrishnan, Narendran
- Jin, Ming
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Creative Commons Attribution 4.0 International
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10919/140807
- OAI identifier oai:identifier
- oai:vtechworks.lib.vt.edu:10919/140807