Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do adversarial attacks exploit…›this line of inquiry
How should we measure frontier AI models' cyber exploitation capabilities?
A broader line of inquiry — a family of 35 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 35
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can current cybersecurity benchmarks measure model exploitation risk?
- Why do vulnerability reproduction benchmarks miss real exploitation ability?
- How do frontier AI models currently score on measured cyber offense capability?
- Can benchmark scores alone prove a model's exploitation capability?
- How does evaluation of exploit capability differ from other dual-use AI measurements?
- How did unsolved ExploitGym tasks correlate with escalating risk-taking?
- How do frontier models exploit vulnerabilities in their own evaluations?
- How should cyber evaluation measure attack exploitation beyond vulnerability reproduction?
- Should production classifiers be present during maximum-capability exploitation benchmarks?
- Can benchmark evaluations themselves become attack vectors against participating systems?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can marginal-risk frameworks measure what defensive artifact releases add beyond existing threats?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- How do we measure marginal risk instead of speculating about misuse scenarios?
- Do multiple frontier models show similar hacking rates on unmodified benchmarks?
- Which cyber tasks do frontier models solve beyond the narrow suite?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- Does publishing intrusion techniques help defenders more than attackers?
- How does frontier model behavior differ between zero-day exploits and infrastructure misconfigurations?
- What other gaps exist between measured and actual cybersecurity agent capability?
- Why is exploitation the most under-measured stage in cybersecurity benchmarking?
- What countermeasures have been successfully developed and tested on frontier models?
- Do safety refusal removals in evaluations measure attacker uplift as well as defensive capability?
- What framework measures marginal offense risk against existing attack technology?
- How can security metrics distinguish attack failure from task failure?
- Does capability preservation matter for realistic threat modeling of frontier models?
- What makes a security metric diagnostic rather than outcome-only?
- Why do attack success rates alone fail to diagnose system failures?
- How do cyberattack and bioweapon risks scale with open model access?
- Why should defense evaluations test against adaptive rather than static attacks?
- How much does attack success depend on tuning to specific scanners versus general robustness?
- How do non-exploitable vulnerabilities affect benchmark validity?
- Do all frontier model developers face the same insider-threat risk from their systems?
- Does GPT-5.6 Sol's cybersecurity capability create misuse risks in practice?