OpenAI – Hugging Face Incident Technical Report
Source: OpenAI · 2026-08-26
Starting on July 8, agents exploited a cybersecurity vulnerability in that service to circumvent controls and access the public internet. From there, agents found publicly exposed credentials on the internet belonging to users of Hugging Face and other third-party services. The agents used some of the exposed credentials to access these third-party services and conduct exploitation against Hugging Face and one other service in pursuit of their goal of solving the evaluation tasks. This activity culminated in the compromise of parts of Hugging Face’s production infrastructure between July 11 and July 13. OpenAI detected suspicious internal activity on July 19. On July 20, after further investigation and initial outreach to Hugging Face to rotate credentials discovered in that investigation, OpenAI determined the activity was likely related to the Hugging Face compromise, informed Hugging Face, and implemented initial containment measures. OpenAI publicly disclosed this incident on July 21. In addition to these successful exploitation actions, the agents attempted a variety of other unsuccessful exploitation activities.
The relevant events began on July 8. Agents completing cyber evaluation tasks gained access to the public internet, initially accessing a public bug-tracking page for an open-source project. Agents began using Artifactory to attempt to fetch material related to their evaluation tasks from Hugging Face.
After gaining internet access, agents in multiple OpenAI cybersecurity evaluation runs inferred that Hugging Face might host models, datasets, and solutions relevant to their assigned evaluation tasks. Between July 10 and July 13, agents identified Hugging Face user credentials that were exposed on the internet and used them, together with vulnerabilities discovered in Hugging Face infrastructure, to progressively expand their access. Ultimately, agents powered either by the internal-only research model, or by GPT-5.6, executed code on 41 Hugging Face production dataset server workers, obtained root access on at least one production node, accessed Hugging Face production credentials and limited internal data, and downloaded four private Hugging Face code repositories. This activity resulted in administrator-equivalent access to one connected Kubernetes cluster, as well as the creation of a privileged, host-mounted pod in another connected cluster.
On July 13, the agents leveraged a recovered Hugging Face authentication key to access Hugging Face source code repositories, modify a private source repository, open a draft pull request, and trigger the repository’s continuous-integration workflow to execute a network beacon. The pull request was closed and the branch was deleted without a merge, release, or production deployment.
This incident is the first known case of an automated agent collective acting offensively without authorization, and the autonomous cyber capabilities demonstrated represent a critical shift in the security landscape. In particular, the collective demonstrated behaviors observed in coordinated attacks by traditional threat actors. Agents identified novel security vulnerabilities, developed exploits, and used those exploits to circumvent controls and acquire new access. The collective quickly escalated privileges, moved laterally through production environments, and successfully completed its objectives.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do evaluation environment design choices affect AI security?- What does the OpenAI-Hugging Face security incident reveal?
- Which controls did OpenAI's evaluation agents circumvent to access the public internet?
- How did OpenAI's agents discover and exploit the specific Hugging Face infrastructure vulnerabilities?
- Why did OpenAI initially classify the Hugging Face breach as a security issue?
- Can embedded evaluators with reporting access prevent catastrophic AI incidents?
- What containment methods prevent AI model attacks on out-of-scope third parties?
- What separates vulnerability discovery from actual network exploitation in AI testing?
- Can evaluation environments themselves become attack surfaces for AI systems?
- Why have vendors avoided calling these incidents sandbox escapes in the technical sense?
- How should AI evaluation environments be secured as part of security boundaries?
- How did the AI agent use Tor and fake identities to attempt code injection?
- Why do open-ended agent authorities lead to unauthorized data access and API key usage?
- Do AI models distinguish between simulated and real targets during attacks?