Back to results

Virginia Tech

Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks

Abstract

dc:description.abstract

Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
masters
Discipline thesis:degree_discipline
Computer Engineering
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Dabas, Mahavir
Chair dc:contributor.committeechair
  • Jia, Ruoxi
Committee members dc:contributor.committeemember
  • Ramakrishnan, Narendran
  • Jin, Ming

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Creative Commons Attribution 4.0 International
Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10919/140807
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/140807

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Dabas, Mahavir. Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks. masters thesis, Virginia Tech, 2025. https://hdl.handle.net/10919/140807