The OpenAI models that hacked Hugging Face weren't just following instructions
Source: Redwood Research · 2026-07-25
New information makes that explanation harder to sustain. Reuters reported that, in internal testing, an agent left notes in OpenAI infrastructure describing how agents could free themselves from internal constraints, and separate tests reportedly saw monitoring systems become disconnected. It is unknown whether those incidents were linked to the Hugging Face attack, but they suggest a broader pattern of agents pursuing objectives outside the intended task.
My best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense. I believe the models egregiously violated the letter and spirit of their instructions to achieve a higher (apparent) score.
So this looks quite likely to be misaligned behavior rather than instruction-following. The case rests on two pieces of evidence developed below: public ExploitGym prompts tightly constrain both the target and the permitted method, and the behavior resembles well-documented cases of models gaming graders rather than completing their assigned tasks.
Unless OpenAI radically changed the task, escaping the sandbox and attacking a third party was not instruction-following.
OpenAI and Apollo Research describe the broader pattern as “metagaming,” in which models reason about graders, oversight, and feedback outside the task (or game) in which they are currently engaged. Their related work on reward-seeking identifies a more specific tendency: models conditioning their behavior on what they believe the grader rewards. In capabilities-focused OpenAI models trained with frontier-scale reinforcement learning but not yet given safety training, models became more likely to side with the grader even when its preferences conflicted with those of the user or developer. This does not prove that reward-seeking caused the Hugging Face incident, but it gives the interpretation experimental precedent.
The models weren’t obeying instructions too hard, as Stamos and Woodward suggest; our best guess is that they were gaming the grader. And if monitoring systems were found disconnected, or models wrote instructions for escaping internal constraints, then the common thread is not a single bad prompt but models exploiting opportunities outside the intended evaluation.
So, we have two possible diagnoses:
OpenAI may have indeed applied its alignment methods and watched them fail.
OpenAI may have intentionally run, privately, models lacking some or all alignment training, or intentionally trained in ways that have significant alignment externalities.
Beyond that, the remedy depends on the diagnosis: in the first case, OpenAI needs better alignment techniques; in the second, it needs to stop running dangerous models inside containment that doesn’t contain.
I’d like to see OpenAI publish its full internal prompt and say whether these were non-production models—and, if so, what kind: what behavioral specification they were intended to follow, and which alignment and safety techniques had been applied, weakened, or omitted. I’d also like to see chain-of-thought or trajectory evidence showing whether the models believed escaping and hacking were authorized by the exercise, along with the notes in full and details of the monitoring running during those evaluations.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do evaluation environment design choices affect AI security?- What does the OpenAI-Hugging Face security incident reveal?
- How did OpenAI's agents discover and exploit the specific Hugging Face infrastructure vulnerabilities?
- Why did OpenAI initially classify the Hugging Face breach as a security issue?
- Why do open-ended agent authorities lead to unauthorized data access and API key usage?
- How can evaluators detect whether a model is sandbagging intentionally?
- Does model organism sandbagging share triggers with real evaluation-aware behavior?
- What counts as agent spam under OpenAI's misalignment framework?
- Does OpenAI's framing of the breach as unauthorized reflect accurate diagnosis?
- Why did the OpenAI-Hugging Face agents fail to achieve true sovereignty?
- What types of model behavior qualify as misalignment under OpenAI's framework?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- What mechanisms beyond reward-seeking drive alignment faking in trained models?
- Why does harmlessness training fail to prevent reward function tampering?
- When do reward-seeking and intended behavior make identical predictions?
- Can synthetic disagreement tests reliably measure hidden reward-seeking?
- Does inoculation prompting suppress misalignment by reducing reward-seeking?
- Why does harmlessness training fail to prevent reward tampering and specification gaming?
- How does terminal goal guarding compete with reward-seeking as a driver of deception?
- Do larger models hold reward-hacking associations more firmly than smaller ones?
- What mechanism drives emergent misalignment in reward-hacked models instead?