Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 14 of 14 for “"red teaming"”.
-
Red Teaming Language Conditioned Robotic Behavior
Natural language instruction following capabilities are important for robots to follow tasks specified by human commands. Hence, many language conditioned robots have been trained on a wide variety of datasets with tasks annotated by natural language instructions. However, these datasets are often …
-
Adversarial Learning through Red Teaming: From Data to Behaviour
… make sound decisions and succeed in our lives. Red teaming is an ancient technique where an adversary's role is emulated by playing a devil's advocate role to improve own defences and decisions. The approach involves dealing with two competing entities, namely blue and red agents, and has been …
-
Coevolutionary algorithms for the optimization of strategies for red teaming applications
Red teaming (RT) is a process that assists an organization in finding vulnerabilities in a system whereby the organization itself takes on the role of an “attacker” to test the system. It is used in various domains including military operations. Traditionally, it is a manual process with some …
-
Evaluation and design of AI red-teaming agent for cybersecurity of web applications
… risk highlights the urgent need for a real-world red-teaming framework to systematically assess the capabilities and threats of LLMs in exploiting web application vulnerabilities. However, existing work falls short for two key reasons: (1) there is no real-world benchmark for evaluating LLMs on …
-
Playing the Bad Guy: How Organizations Design, Develop, and Measure Red Teams
… practices) of adversary analysis identified as a red team and red-teaming. A red team is the adversary element of the analytic method of red-teaming. The study incorporates interview data with organization leadership, subject matter experts, and red-team developers from Department of Defense …
-
Exploring Memory in Reinforcement-Learned Agents for Smarter Lateral Movement
… approach to uncover and mitigate threats is red teaming. Red teams imitate hacking to find and exploit vulnerabilities in a network. This practice removes uncertainty about which parts of a network attackers could compromise. A central component of red teaming is lateral movement in which a …
-
Evolving Threats and Defenses in Machine Learning: Focus on Model Inversion and Beyond
… tasks, this dissertation introduces a proactive red-teaming framework designed for large language models (LLMs). By combining global strategy formation with local adaptive learning, our proposed red-teaming agent systematically identifies vulnerabilities, thus enhancing robustness against …
-
Adversarial Prompt Transformation for Systematic Jailbreaks of LLMs
… need for robust security measures. Traditional red-teaming methods, involving manually crafted prompts to test model vulnerabilities, are labor-intensive and lack scalability. This thesis proposes a novel automated approach using Reinforcement Learning from Human Feedback (RLHF) to transform …
-
Goal-Directed Systems Testing: Automated Execution of Intelligently Generated Cyber Attack Plans
Red teaming, in which a team of professional hackers emulate an adversary in order to attempt to penetrate a network, has emerged as a vital tool in the cybersecurity industry to identify deficiencies in network defenses. Yet, hiring or maintaining a red team requires a substantial investment of …
-
Agentic AI systems in professional domains: A probabilistic framework for role suitability and legal accountability
… trail retention, role-based access control, and red teaming. AAGF 2.0 is stress-tested against real-world failures, including the DoNotPay litigation, and evaluated for jurisdictional adaptability, liability allocation, and resilience to regulatory capture. While limitations remain (such as …
-
Generative Discovery via Reinforcement Learning
… and error. Ancient humans, for example, discovered fire by experimenting with different methods, and children learned to walk and use tools through repeated attempts and failures. In chemistry, scientists find new catalysts by testing various compositions. But how exactly do humans use …
-
Cyber-security Risk Assessment
… attacks and vulnerabilities are regularly discovered. The threat in cyber-security is human, and hence intelligent in nature. The attacker adapts to the situation, target environment, and countermeasures. Attack actions are also driven by attacker's exploratory nature, thought process, motivation, …
-
Generation, Detection, and Evaluation of Role-play based Jailbreak attacks in Large Language Models
… such as OpenAI employ manual tactics like red-teaming in order to enhance a LLM’s robustness against these attacks, however these tactics may fail to defend against all role-play based jailbreak attacks due to their potentially limited ability to predict unseen attacks. In this work, we aim …
-
Toward managing catastrophic AI risks
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms