Path to Astra: critical capabilities and frontier safeguards
Source: OpenAI · 2026-09-01
Since our earlier assessment that Astra might reach a critical level of cybersecurity capability, we have gathered more evidence and run additional evaluations to assess the model’s capabilities. We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. It is the first model we are designating at this level, and requires stronger safeguards during development and before release.
Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions. Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.
We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use.
The model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.
The model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
Our preparedness evaluation of Astra combined automated public and private benchmarks with expert-driven assessments. Astra represents a significant increase in cybersecurity capabilities compared to GPT‐5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development.
As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
In expert-led assessments against a hardened browser and operating system, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains. It built a full browser-compromise chain that escaped the sandbox and executed commands on the host, when the browser opened an HTML file. The model also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root. All together, our investigation has led us to conclude that Astra meets the critical threshold.
For models with Astra’s level of cybersecurity capabilities, we need to cover two pathways to minimize risk for severe cyber harm, both during development and before deployment:
Notably, the second pathway applies to both internal development and external deployment. As we previously described, we paused certain frontier training (including certain training for Astra) for two weeks after the OpenAI-Hugging Face incident in order to harden our training infrastructure, including isolation and network controls, expanded monitoring, and strengthened alignment training and thresholds. We then continued smaller-scale work under stricter controls.
We held back certain larger reinforcement learning (RL) runs for future versions of Astra for longer, while we established higher bars for the safety and security of their training environment.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits recursive self-improvement in autonomous AI systems? How should we measure frontier AI models' cyber exploitation capabilities?- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- What other gaps exist between measured and actual cybersecurity agent capability?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- What framework measures marginal offense risk against existing attack technology?
- How do non-exploitable vulnerabilities affect benchmark validity?
- What countermeasures have been successfully developed and tested on frontier models?
- How can security metrics distinguish attack failure from task failure?
- What makes a security metric diagnostic rather than outcome-only?
- Why do attack success rates alone fail to diagnose system failures?
- How do cyberattack and bioweapon risks scale with open model access?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- Can marginal-risk frameworks measure what defensive artifact releases add beyond existing threats?
- Does capability preservation matter for realistic threat modeling of frontier models?
- How do benchmark scores differ from deployment safety requirements?
- How do single-axis safety benchmarks misrepresent deployment readiness?