Pacing model development in an era of cyber-critical capabilities

Paper · Source
Frontier AI Risk & RSI

Source: OpenAI · 2026-08-18

Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework. Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process.

Alignment—the work of making AI systems behave as intended and responsive to human oversight—has long been at the core of our research program. We now require stronger evidence of aligned behavior throughout all of training, building on research and evaluations already underway. Keeping increasingly capable systems aligned is a challenge the whole field will need to address. The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework.

Our approach to developing more capable models rests on three reinforcing safeguards:

Monitoring, which detects and allows us to respond to concerning behavior.

Alignment, which reduces the likelihood of harmful or unauthorized actions.

Security measures, which limit what AI systems can access or affect.

We expect models to soon drive most security work, including defending against other models. This will allow all three safeguards to scale with model capability, which we see as crucial.

Workload isolation: We now require stronger isolation (“sandboxes”) for workloads that execute model-generated or otherwise untrusted code. This also applies to software that could be compromised while processing model outputs.

Network isolation: We have implemented more controls to isolate higher-risk and untrusted workloads from the internet. We have designed these controls so that a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the internet or other internal networks.

Continuous security testing: We have reconfigured our environment to remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries. We are also improving our ability to collect and monitor security logs. Finally, we are investing in automation using our models to test these boundaries continuously against simulated attacks.

Together, these controls provide defense in depth and help protect our research environment against attacks. Today, we require the strictest level of security safeguards for workloads involving Astra or cyber models, since we have determined that Astra models may have a critical level of cyber capability. These safeguards also apply to all other cyber-related workloads.

While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar. We are prioritizing safety and alignment workloads for migration to these new environments first.

This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation environment design choices affect AI security? How does awareness of evaluation context influence model behavior? How do individually-safe actions create collectively-unsafe outcomes? What governance mechanisms can effectively constrain widely deployed AI systems? Do individually safe AI actions create unsafe outcomes in integrated systems? Why do standard evaluation practices obscure safety-critical AI failures? How can humans maintain effective oversight as AI systems scale? Can AI research automation sustain progress through accelerating feedback loops? What limits recursive self-improvement in autonomous AI systems?