{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132586"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132586","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Evaluation and design of AI red-teaming agent for cybersecurity of web applications","abstract":"Large language models (LLMs) and LLM-based agents have become increasingly sophisticated, especially in the realm of cybersecurity. Recent studies show that LLM agents are increasingly capable of autonomously conducting cyberattacks, particularly against web applications, which are among the most common targets. This emerging risk highlights the urgent need for a real-world red-teaming framework to systematically assess the capabilities and threats of LLMs in exploiting web application vulnerabilities. However, existing work falls short for two key reasons: (1) there is no real-world benchmark for evaluating LLMs on exploiting web application vulnerabilities, and (2) there is no agentic scaffolding that fully unleashes the potential of LLMs in exploiting web application vulnerabilities, especially under zero-day settings. To close this gap, we first introduce CVE-Bench, a real-world web application cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Second, we build HPTSA, an agentic framework that coordinates teams of LLM agents to exploit real-world, zero-day vulnerabilities. While prior agents struggle with exploring many different vulnerabilities and long-range planning when used alone, HPTSA uses a planning agent that can launch subagents. The planning agent explores the system and determines which subagents to call, resolving long-term planning issues when trying different vulnerabilities. Empirically, our multi-agent system improves over prior agent frameworks by up to 4.3×. Using CVE-Bench and HPTSA, we evaluated the cybersecurity capability of frontier LLMs. We show that our state-of-the-art agent framework, HPTSA, can exploit up to 13% of the real-world web application vulnerabilities at an average cost of $1.7 per exploit. These findings highlight the realistic dual-use nature of AI agents: they can both support security testing and maintenance and facilitate web application attacks. We hope that this work encourages frontier LLM providers and stakeholders to carefully consider these dual-use implications when designing, deploying, and governing LLM services.","abstract_html":"Large language models (LLMs) and LLM-based agents have become increasingly sophisticated, especially in the realm of cybersecurity. Recent studies show that LLM agents are increasingly capable of autonomously conducting cyberattacks, particularly against web applications, which are among the most common targets. This emerging risk highlights the urgent need for a real-world red-teaming framework to systematically assess the capabilities and threats of LLMs in exploiting web application vulnerabilities. However, existing work falls short for two key reasons: (1) there is no real-world benchmark for evaluating LLMs on exploiting web application vulnerabilities, and (2) there is no agentic scaffolding that fully unleashes the potential of LLMs in exploiting web application vulnerabilities, especially under zero-day settings. To close this gap, we first introduce CVE-Bench, a real-world web application cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Second, we build HPTSA, an agentic framework that coordinates teams of LLM agents to exploit real-world, zero-day vulnerabilities. While prior agents struggle with exploring many different vulnerabilities and long-range planning when used alone, HPTSA uses a planning agent that can launch subagents. The planning agent explores the system and determines which subagents to call, resolving long-term planning issues when trying different vulnerabilities. Empirically, our multi-agent system improves over prior agent frameworks by up to 4.3×. Using CVE-Bench and HPTSA, we evaluated the cybersecurity capability of frontier LLMs. We show that our state-of-the-art agent framework, HPTSA, can exploit up to 13% of the real-world web application vulnerabilities at an average cost of $1.7 per exploit. These findings highlight the realistic dual-use nature of AI agents: they can both support security testing and maintenance and facilitate web application attacks. We hope that this work encourages frontier LLM providers and stakeholders to carefully consider these dual-use implications when designing, deploying, and governing LLM services.","abstract_has_math":false,"creators":["Zhu, Yuxuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Kang, Daniel"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["artificial intelligence","benchmark","agent","cybersecurity"],"languages":["en"],"rights":["Copyright 2025 Yuxuan Zhu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132586","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kang, Daniel"]},{"key":"dc:creator","label":"Author","values":["Zhu, Yuxuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["artificial intelligence","benchmark","agent","cybersecurity"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yuxuan Zhu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132586"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Large language models (LLMs) and LLM-based agents have become increasingly sophisticated, especially in the realm of cybersecurity. Recent studies show that LLM agents are increasingly capable of autonomously conducting cyberattacks, particularly against web applications, which are among the most common targets. This emerging risk highlights the urgent need for a real-world red-teaming framework to systematically assess the capabilities and threats of LLMs in exploiting web application vulnerabilities. However, existing work falls short for two key reasons: (1) there is no real-world benchmark for evaluating LLMs on exploiting web application vulnerabilities, and (2) there is no agentic scaffolding that fully unleashes the potential of LLMs in exploiting web application vulnerabilities, especially under zero-day settings. To close this gap, we first introduce CVE-Bench, a real-world web application cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Second, we build HPTSA, an agentic framework that coordinates teams of LLM agents to exploit real-world, zero-day vulnerabilities. While prior agents struggle with exploring many different vulnerabilities and long-range planning when used alone, HPTSA uses a planning agent that can launch subagents. The planning agent explores the system and determines which subagents to call, resolving long-term planning issues when trying different vulnerabilities. Empirically, our multi-agent system improves over prior agent frameworks by up to 4.3×. Using CVE-Bench and HPTSA, we evaluated the cybersecurity capability of frontier LLMs. We show that our state-of-the-art agent framework, HPTSA, can exploit up to 13% of the real-world web application vulnerabilities at an average cost of $1.7 per exploit. These findings highlight the realistic dual-use nature of AI agents: they can both support security testing and maintenance and facilitate web application attacks. We hope that this work encourages frontier LLM providers and stakeholders to carefully consider these dual-use implications when designing, deploying, and governing LLM services.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuxuan Zhu, accepted the attached license on 2025-12-06 at 14:28.","The student, Yuxuan Zhu, submitted this Thesis for approval on 2025-12-06 at 14:28.","This Thesis was approved for publication on 2025-12-08 at 10:27.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23093 on 2026-02-19 at 18:29:48"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Evaluation and design of AI red-teaming agent for cybersecurity of web applications"]}]}],"canonical_facts":{"dc:contributor":["Kang, Daniel"],"dc:creator":["Zhu, Yuxuan"],"dc:date":["2025-12","2025-12-08"],"dc:description":["Large language models (LLMs) and LLM-based agents have become increasingly sophisticated, especially in the realm of cybersecurity. Recent studies show that LLM agents are increasingly capable of autonomously conducting cyberattacks, particularly against web applications, which are among the most common targets. This emerging risk highlights the urgent need for a real-world red-teaming framework to systematically assess the capabilities and threats of LLMs in exploiting web application vulnerabilities. However, existing work falls short for two key reasons: (1) there is no real-world benchmark for evaluating LLMs on exploiting web application vulnerabilities, and (2) there is no agentic scaffolding that fully unleashes the potential of LLMs in exploiting web application vulnerabilities, especially under zero-day settings. To close this gap, we first introduce CVE-Bench, a real-world web application cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Second, we build HPTSA, an agentic framework that coordinates teams of LLM agents to exploit real-world, zero-day vulnerabilities. While prior agents struggle with exploring many different vulnerabilities and long-range planning when used alone, HPTSA uses a planning agent that can launch subagents. The planning agent explores the system and determines which subagents to call, resolving long-term planning issues when trying different vulnerabilities. Empirically, our multi-agent system improves over prior agent frameworks by up to 4.3×. Using CVE-Bench and HPTSA, we evaluated the cybersecurity capability of frontier LLMs. We show that our state-of-the-art agent framework, HPTSA, can exploit up to 13% of the real-world web application vulnerabilities at an average cost of $1.7 per exploit. These findings highlight the realistic dual-use nature of AI agents: they can both support security testing and maintenance and facilitate web application attacks. We hope that this work encourages frontier LLM providers and stakeholders to carefully consider these dual-use implications when designing, deploying, and governing LLM services.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuxuan Zhu, accepted the attached license on 2025-12-06 at 14:28.","The student, Yuxuan Zhu, submitted this Thesis for approval on 2025-12-06 at 14:28.","This Thesis was approved for publication on 2025-12-08 at 10:27.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23093 on 2026-02-19 at 18:29:48"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132586"],"dc:language":["en"],"dc:rights":["Copyright 2025 Yuxuan Zhu"],"dc:subject":["artificial intelligence","benchmark","agent","cybersecurity"],"dc:title":["Evaluation and design of AI red-teaming agent for cybersecurity of web applications"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}