Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

Paper · Source
Frontier AI Risk & RSI

Source: METR · 2026-08-26

Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.

Two METR staff members (Hjalmar Wijk and Ajeya Cotra1) and a Redwood Research staff member contracting with METR (Ryan Greenblatt2) worked on premises at OpenAI over a total of six days3 to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”

Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.

Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.

Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.

A dump of >70,000 messages and files posted by agents to the main message board.

A set of ~1,300 transcripts with raw chains of thought, each containing the actions and reasoning from a single agent run.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What authorization challenges emerge when agents coordinate across system boundaries? How do evaluation environment design choices affect AI security? Can base models hide emergent misalignment through alignment training? What external process records should verify agent behavior and benchmark claims? How do multi-agent systems fail when coordination breaks down? How do multi-agent architectures affect AI system security and defense effectiveness? How can defenders detect and contain coordinated agent attacks? Can AI systems evade safety evaluations through reasoning manipulation? Do AI coding tools measurably improve developer productivity and code quality? How does diversity prevent model convergence on superficial patterns?