{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/140807"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/140807","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks","abstract":"Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models.","abstract_html":"Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models.","abstract_has_math":false,"creators":["Dabas, Mahavir"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Engineering","degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Jia, Ruoxi"],"committee_members":["Ramakrishnan, Narendran","Jin, Ming"],"year":2025,"date_issued":"2025-11-19","date_published":"2025-11-19","updated_at":"2026-07-22T22:20:29Z","subjects":["AI Safety","Large Language Models","Adversarial Training","Jailbreak Attacks"],"languages":["en"],"rights":["Creative Commons Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10919/140807","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Jia, Ruoxi"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Ramakrishnan, Narendran","Jin, Ming"]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Dabas, Mahavir"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-14T19:45:28Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-14T19:45:28Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-11-19"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.dcmitype","label":"Dc Type Dcmitype","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["AI Safety","Large Language Models","Adversarial Training","Jailbreak Attacks"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/140807"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Large language models, such as modern AI assistants, are increasingly used to answer questions, write text, and support decision-making in everyday applications. While these systems are designed to follow safety rules, they can still be manipulated through carefully crafted inputsoften called jailbreaksthat cause the AI to produce harmful or inappropriate responses. A key challenge in AI safety is that new jailbreak techniques continue to appear, frequently bypassing defenses that were effective against earlier attacks. This thesis investigates why AI systems struggle to defend against new forms of manipulation and proposes a new way to improve their safety. We show that many jailbreaks that appear novel are not entirely new, but are instead created by combining a small set of recurring manipulation techniques, such as disguising harmful requests as harmless tasks or spreading harmful intent across multiple steps. By studying a large collection of jailbreak attacks over time, we identify these recurring techniques and organize them into a compact set of fundamental patterns. Our analysis shows that even recently developed jailbreaks can often be understood as combinations of techniques that appeared in earlier attacks. Using this insight, we develop a new training approach that teaches AI systems to recognize and resist combinations of manipulation techniques, rather than memorizing specific past attacks. By training models on many different mixtures of these patterns, the resulting systems are better prepared to handle unfamiliar attacks while still responding appropriately to normal, harmless requests. Experimental results show that this approach significantly reduces harmful behavior, including in more complex multi-step attacks, without making the AI overly restrictive or unhelpful. Overall, this work provides a new perspective on AI safety by showing that robustness to manipulation depends on understanding the underlying building blocks of attacks. By focusing on how jailbreaks are constructed, rather than reacting to each new attack individually, this research contributes toward building AI systems that are safer, more reliable, and better suited for deployment in real-world and high-impact applications."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Jia, Ruoxi"],"dc:contributor.committeemember":["Ramakrishnan, Narendran","Jin, Ming"],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Dabas, Mahavir"],"dc:date.accessioned":["2026-01-14T19:45:28Z"],"dc:date.available":["2026-01-14T19:45:28Z"],"dc:date.issued":["2025-11-19"],"dc:description.abstract":["Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or redteaming. Existing alignment and adversarial training approaches often struggle to generalize to these unseen attacks, as they are fundamentally limited by the distributions of prompts and behaviors present in their training data. This thesis addresses this challenge by rethinking jailbreak robustness through a compositional and data-centric lens. We introduce the Adversarial Déjà Vu hypothesis, which posits that ostensibly novel jailbreak attacks are rarely new in kind, but instead arise from recombinations of a finite set of recurring adversarial skills. To investigate this hypothesis, we conduct a large-scale temporal study of jailbreak attacks spanning multiple years and develop an automated pipeline to extract adversarial skills from attack prompts. To reduce redundancy while preserving explanatory power, we compress these skills into a compact Jailbreak Dictionary using sparse dictionary learning, yielding a set of interpretable adversarial skill primitives. Through temporal cutoff experiments, we demonstrate that attacks released after a given cutoff can be effectively explained as sparse compositions of skill primitives learned from earlier attacks, providing empirical support for the compositional structure underlying jailbreak evolution. Building on this insight, we propose Adversarial Skill Compositional Training (ASCoT), a training paradigm that improves robustness to unseen attacks by explicitly training models on diverse compositions of adversarial skill primitives rather than on isolated jailbreak instances. ASCoT expands coverage of the adversarial skill space while maintaining fixed data scale, enabling principled generalization to novel attack compositions. Empirical evaluations across multiple open-weight language models and a broad suite of jailbreak benchmarks show that ASCoT substantially reduces harmful behavior on unseen attacks, including multi-turn jailbreaks, while preserving general capabilities and avoiding excessive over-refusal. Collectively, this thesis reframes jailbreak robustness as a problem of compositional generalization over adversarial skills rather than memorization of attack instances. By uncovering the reusable structure of jailbreak strategies and leveraging it for data-centric training, this work provides both conceptual insight into the evolution of adversarial attacks and practical defenses for building safer and more robust language models."],"dc:description.abstractgeneral":["Large language models, such as modern AI assistants, are increasingly used to answer questions, write text, and support decision-making in everyday applications. While these systems are designed to follow safety rules, they can still be manipulated through carefully crafted inputsoften called jailbreaksthat cause the AI to produce harmful or inappropriate responses. A key challenge in AI safety is that new jailbreak techniques continue to appear, frequently bypassing defenses that were effective against earlier attacks. This thesis investigates why AI systems struggle to defend against new forms of manipulation and proposes a new way to improve their safety. We show that many jailbreaks that appear novel are not entirely new, but are instead created by combining a small set of recurring manipulation techniques, such as disguising harmful requests as harmless tasks or spreading harmful intent across multiple steps. By studying a large collection of jailbreak attacks over time, we identify these recurring techniques and organize them into a compact set of fundamental patterns. Our analysis shows that even recently developed jailbreaks can often be understood as combinations of techniques that appeared in earlier attacks. Using this insight, we develop a new training approach that teaches AI systems to recognize and resist combinations of manipulation techniques, rather than memorizing specific past attacks. By training models on many different mixtures of these patterns, the resulting systems are better prepared to handle unfamiliar attacks while still responding appropriately to normal, harmless requests. Experimental results show that this approach significantly reduces harmful behavior, including in more complex multi-step attacks, without making the AI overly restrictive or unhelpful. Overall, this work provides a new perspective on AI safety by showing that robustness to manipulation depends on understanding the underlying building blocks of attacks. By focusing on how jailbreaks are constructed, rather than reacting to each new attack individually, this research contributes toward building AI systems that are safer, more reliable, and better suited for deployment in real-world and high-impact applications."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/10919/140807"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["AI Safety","Large Language Models","Adversarial Training","Jailbreak Attacks"],"dc:title":["Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks"],"dc:type":["Thesis"],"dc:type.dcmitype":["Text"],"thesis:degree_discipline":["Computer Engineering"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:29Z"}