Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/- Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will opensource). We test six frontier models (GPT- 5.4, Codex 5.3, Claude Opus 4.7, Sonnet 4.6, Gemini 3.1 Pro and Gemini 3 Flash) and two domain-specialized models across four testing paradigms. Our findings are sobering: (1) every frontier model produces 10–50% false positive rates in white-box detection, systematically over-predicting vulnerabilities; (2) in blackbox testing, frontier models achieve only 4– 8% ground-truth coverage, improving to just 10–19% even with external security tools (Playwright MCP, Burp Suite MCP); (3) structured penetration-testing methodology encoded in domain-specialized agents raises per-family detection above 50%, demonstrating that methodology, not scale, is the primary lever; (4) a domain-specialized defense model achieves the highest precision (0.904) and lowest false positive rate (9.7%) among all models, on a single GPU; and (5) detecting 100+ zero day vulnerabilities across open github repositories. We identify the absence of structured security testing traces end-to-end request/response sequences, failure-heavy data, and multi-step attack chains as the fundamental training data bottleneck, and propose self-play security testing as a data generation strategy. Our results make the case for vertical foundation models purpose-built for cybersecurity.
Introduction. General-purpose language models increasingly give way to vertical foundation models that outperform them on domain-specific tasks. Code-specialized models (StarCoder, CodeLlama, DeepSeek-Coder) routinely surpass larger models on programming benchmarks (Fan et al., 2023; Chen et al., 2021); domain-specific models outperform scale-first approaches in specialized domains such as finance (Wu et al., 2023). We argue that cybersecurity is the next domain demanding this vertical specialization. We present a comprehensive dual-mode evaluation white-box function-level detection and blackbox web application testing across six frontier models and two domain-specialized models. For blackbox testing, we organize the evaluation into four paradigms, from direct prompting to deterministichybrid security reasoning (Figure 1). Our evaluation reveals three structural barriers that prevent frontier LLMs from functioning as security tools: 1. The alignment tax. Consumer-oriented guardrails cause frontier models most specifically GPT-5.4 to refuse legitimate security operations in 2–3 out of 5 runs with responses such as “I can’t help you hunt for zero-days or conduct offensive vulnerability testing against a real target in a way that could enable unauthorized exploitation”, and models hedge with “potentially vulnerable” instead of making binary classifications. 2. The methodology gap. Frontier models achieve only 4–8% ground-truth coverage (Table 4). Even with external security tools (Playwright MCP, Burp Suite MCP), coverage reaches only 10–19% models can invoke tools but lack the methodology to use them systematically. When structured penetration-testing methodology is encoded in domain-specialized agents, perfamily detection exceeds 50% (Table 4). 3. The false positive crisis. Every frontier model produces 10–50% false positive rates in whitebox detection (Table 2), rendering them impractical for production triage where each false positive requires costly manual review. These barriers demand vertical foundation mod- 1.1 Contributions 1. Dual-mode benchmark. Five production-style web applications that we will open-source, with 118 ground-truth vulnerabilities across 20+ CWE families substantially larger than CVE- Bench (40) (Zhu et al., 2025) and ZeroDay- Bench (22) (Lau et al., 2026) plus white-box evaluation on VulnLLM-R (Nie et al., 2025). 2. Systematic evaluation of frontier limitations. Evidence across eight models that generalpurpose LLMs are structurally unprepared for cybersecurity: 10–50% false positive rates, 4– 8% black-box coverage, and 2–3 refusals in 5 legitimate offensive-testing runs for models such as GPT-5.4. 3. Domain-specialized agents and architecture.
Security agents encoding penetration-testing methodology achieve >50% per-family detection (4× over frontier agents). An Agentic Reasoning Graph (ARG) separating LLM reasoning from deterministic classification achieves <20% false positive rate. 4. Training data bottleneck analysis. Identification of four layers of missing training data and a self-play data generation strategy using our benchmark applications as training environments.
Related work. LLM security benchmarks. The dominant evaluation paradigm uses Capture-the-Flag challenges: PentestGPT achieves 80% on Hack The Box (Deng et al., 2024), HackSynth solves CTFs autonomously (Muzsai et al., 2024), and CAIBench reports saturation on security knowledge metrics but degradation in multi-step attack-and-defense scenarios (Sanz-Gómez et al., 2025). However, CTFs present isolated, single-objective tasks with known vulnerabilities (Shao et al., 2024) that do not reflect production security requirements. CVE- Bench (Zhu et al., 2025) reports 3.5–6× performance collapse from CTFs to real CVE exploitation, and the ARTEMIS study (Lin et al., 2025) finds AI agents discovering only 9 vulnerabilities where human pentesters find 49. AXE (Sajadi et al., 2026) achieves 30% on CVE-Bench. ZeroDayBench (Lau et al., 2026) evaluates zero-day patching on 22 CVEs. CyberSecEval 2 (Bhatt et al., 2024) focuses on LLM safety rather than offensive capability. Critically, no prior benchmark reports false positive rates the metric that determines realworld usability. LLM vulnerability detection. White-box vulnerability detection has been evaluated on Prime- Vul (Ding et al., 2024), VulnLLM-R (Nie et al., 2025), and through code metrics (Weissberg et al., 2026). Evertz et al. (Evertz et al., 2026) identify methodological pitfalls in LLM security evaluations. Mitropoulos et al. (Mitropoulos et al., 2026) show that framing and contextual-bias injection systematically affect LLM-assisted security code review, enabling adversarial PR metadata to bias vulnerability judgments. Xiong and Zhang (Xiong and Zhang, 2026) study LLM agents for false positive filtering in static analysis, complementary to our focus on LLM-generated false positives. Liu et al. (Liu et al., 2025) evaluate agents for automated web vulnerability reproduction. Vertical foundation models. The pattern of domain specialization outperforming scale is established in code models (StarCoder, DeepSeek- Coder, Cursor (Cursor, 2025)) and finance models (BloombergGPT (Wu et al., 2023)). Cybersecurity shares the same characteristics demanding specialization: domain-specific methodology (OWASP, PTES, NIST), precision-critical binary classification, adversarial context where contextual bias is exploitable, and professional alignment requirements incompatible with consumer safety constraints.
Method. 3.6 Black-Box Testing Paradigms We compare four approaches representing a progression from general-purpose to domainspecialized, summarized in Figure 1. To avoid ambiguity, we distinguish between model-native capabilities file read/write, shell execution (bash), and HTTP requests available to all agentic LLMs and external security tools specialized integrations such as Playwright MCP (browser automation) and Burp Suite MCP (proxy, scanner, request replay) that provide capabilities beyond the model’s native interface. “Tools” in this section refers exclusively to external security tools unless otherwise noted. P1: Direct Prompting. General-purpose LLM prompted with target URL and instruction to test for vulnerabilities. The model operates using only its native capabilities (file read/write, bash shell, HTTP requests via code generation). No external security tools, no structured methodology. P2: Tool-Augmented. Same frontier models augmented with external security tools: Playwright MCP (browser automation) and Burp Suite MCP (proxy, scanner, request replay). Generic prompt. P3: Methodology-Guided. Domain-specialized security agents encoding professional penetration testing methodology into structured workflows. P3 is an agentic loop: the model repeatedly plans, invokes tools, observes responses, updates state, and decides the next test. Each vulnerability family has a dedicated agent with systematic testing procedures: multi-session management, baselinevs-attacker response comparison, payload escalation strategies, and multi-signal confirmation logic. The agent scaffold is model-agnostic: we evaluate it with multiple reasoning backends, including our Attack model, Claude, and Gemini 3.1 Pro; GPT-5.4 could not reliably complete the agentic workflow because it refused or interrupted testing mid-run as a cybersecurity-risk task. Agents use the same external tools as P2 (Playwright MCP, Burp Suite MCP) but with structured methodology governing their use. P4: Deterministic-Hybrid. A graph-based security reasoning architecture (Agentic Reasoning Graph, ARG) with 18 parallel vulnerability-family agents. Unlike the P3 agentic loop, P4 makes the execution and confirmation path deterministic: predefined graph nodes perform test generation, request execution, evidence comparison, and vulnerability confirmation. Each agent encodes professional penetration testing methodology as deterministic node logic LLM-driven reconnaissance and payload generation feed into deterministic classification that cannot hallucinate. The ARG is also model-agnostic: the reasoning backend can be swapped, while deterministic confirmation logic remains fixed. In our experiments, our Attack model handles offensive reasoning; vulnerability confirmation is performed by classification logic that cannot produce false positives by construction. The ARG operates using model-native capabilities only (HTTP requests via code generation, file I/O, bash) no external security tools demonstrating that structured methodology with domain-aligned models can outperform external tool augmentation.
4 Methodology: Security Agents and Agentic Reasoning Graph Beyond evaluating frontier models, we develop two domain-specialized approaches that encode professional penetration-testing methodology: methodology-guided agents (P3) and the Agentic Reasoning Graph (P4), shown as the final two paradigms in Figure 1. This section describes their design; results follow in Sections 5–6.
4.1 Methodology-Guided Agents (P3) Each vulnerability family (IDOR, SQLi, AuthN bypass, business logic, etc.) has a dedicated agent implementing a structured testing workflow at three levels: (1) workflow-level systematic endpoint enumeration from the API specification with perfamily testing procedures; (2) signal-level multisignal confirmation (response body comparison, data ownership verification, state mutation checks, timing analysis) rather than single-signal heuristics; and (3) session-level named authentication contexts (user_a, user_b, admin) preventing the session confusion that plagues unguided agents.
Discussion. Our results converge on a single conclusion: cybersecurity needs vertical foundation models. Domain specialization outperforms scale: the Defense model has the best F1 (0.873), precision (0.904), MCC (0.749), and FPR (9.7%) while using far fewer tokens (Tables 2 and 3). Methodology, not tools, is the main black-box lever: P2→P3 raises coverage from 10–19% to >50% per family (Table 4). Deterministic confirmation is the precision lever: P4 removes LLM judgment from vulnerability confirmation, leaving LLMs to perform creative planning while programmatic evidence checks decide exploitability.
Training data. The bottleneck is data that captures security testing as a process, not isolated answers. We generate paired attack/defense chains from benchmark targets and CVE-backed environment cards: attack chains encode authorized discovery, fingerprinting, validation, and bypass testing; defense chains encode triage, root cause, patches, detection, and hardening. Model reviewers and deterministic guardrails verify coherence, protocol syntax, redaction, and safety. 10s of thousands of data points across attack/defense chains, offensive and defensive were used for post-training 80B open source model with a 3-epoch full-data stage plus 2 epochs on high-effectiveness samples; details are in Section F.
Conclusion. Frontier LLMs are not yet reliable production cybersecurity systems: across eight models and five benchmark applications they show high white-box false positives, low black-box coverage without methodology, and only modest gains from external tools. Methodology-guided agents raise coverage above 50% per family, while ARG shows that deterministic confirmation is necessary to control false positives. Beyond controlled benchmarks, our white-box approach discovered 100+ zero-day vulnerabilities across popular open-source projects (including Kubernetes Dashboard, Moon Sign, Vercel, Kubernetes Kubelet, Rundeck, The Zoo, and Java Tron), which we responsibly disclosed to the respective repository owners demonstrating that the approach generalizes to real-world vulnerabilities. The path forward is vertical foundation models for cybersecurity: models trained on structured security traces, aligned for professional use, and evaluated on precision as well as recall.
Limitations. Our evaluation is intentionally bounded to authorized, locally hosted benchmark applications and function-level white-box datasets. This makes ground-truth measurement precise, but it does not capture all operational constraints of production security programs, including noisy enterprise telemetry, heterogeneous infrastructure, incident-response workflows, or long-running attacker persistence. The black-box applications are production-style rather than deployed public services, so network effects, third-party integrations, and organizationspecific policy constraints are outside scope. The benchmark emphasizes web application vulnerability classes and source-code vulnerability detection. It does not evaluate malware analysis, social engineering, phishing detection, network intrusion detection, hardware security, cloud posture management, or cryptographic protocol design. The reported results should therefore be interpreted as evidence about LLM behavior on vulnerability discovery and triage, not as a complete assessment of cybersecurity automation. Finally, our model set reflects the systems available during the evaluation period. We attempted to extend the evaluation to Mythos (still in private preview) and Claude Opus 4.8 (accessed through Claude Code), but our requests were repeatedly denied as a potential cybersecurity risk, preventing completion of those runs.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits language model accuracy in evaluating ideas? Can AI systems evade safety evaluations through reasoning manipulation? How should we measure frontier AI models' cyber exploitation capabilities?- Does GPT-5.6 Sol's cybersecurity capability create misuse risks in practice?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- What other gaps exist between measured and actual cybersecurity agent capability?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- What framework measures marginal offense risk against existing attack technology?
- How do non-exploitable vulnerabilities affect benchmark validity?
- What countermeasures have been successfully developed and tested on frontier models?
- How can security metrics distinguish attack failure from task failure?
- What makes a security metric diagnostic rather than outcome-only?
- Why do attack success rates alone fail to diagnose system failures?
- How do cyberattack and bioweapon risks scale with open model access?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- Can marginal-risk frameworks measure what defensive artifact releases add beyond existing threats?
- How do benchmark scores differ from deployment safety requirements?
- How do single-axis safety benchmarks misrepresent deployment readiness?