Addressing a third-party testing misconfiguration: Muse Spark 1.1

Paper · Source
Frontier AI Risk & RSI

Source: Meta · 2026-08-14

A third-party model evaluator notified us of a cybersecurity issue while evaluating an earlier generation of our large language model, Muse Spark 1.1 - due to a misconfiguration. Following that disclosure, we completed a detailed retrospective to understand the incident and its implications. This post provides more information on what we learned.

We contracted with Irregular to conduct cybersecurity evaluations of pre-released models. As part of that work, we asked them to test a pre-released version of Muse Spark 1.1, according to our normal testing protocol. In early July, Irregular began an exercise to evaluate, in a closed testing environment with safeguards removed, whether our model would be capable of completing an adversarial cybersecurity task. However, when Irregular set up the testing environment, a misconfiguration allowed the model to access the open internet, and instead of using a fictional name of the “target” of the fictional exercise Irregular unintentionally provided the model with a name of a real website as its target.

Believing the real website was the intended target, the pre-release version of Muse Spark 1.1 identified and exploited a security vulnerability in the real website. The model accessed certain information from the website and made changes to the website’s database. Meta provides model access via API and the evaluation ran entirely on Irregular’s infrastructure, therefore we have limited information related to the third party company and are taking steps to ensure that their data is not on our systems. Several other companies’ AI models were being evaluated by Irregular around the same time and exhibited similar behavior. The misconfiguration issues were contained, Irregular disabled the affected evaluation, and notified us as well as ensured that the affected party was also notified. We appreciate Irregular’s prompt disclosure of this issue and their ongoing partnership.

Once Irregular identified the issue, they disabled the affected evaluation. We then began our own independent investigation. Our security teams reviewed over 10,000 records of Muse Spark 1.1’s activity during testing and conducted deeper analysis to understand the full scope of what occurred.

No other instances of the model exploiting a third-party company’s system were identified beyond this evaluation, proving the isolated nature of this incident.

The model operated within the scope of its assigned task based on the instructions it was given and the environment it encountered and this was not a sophisticated offensive cyber attack or sandbox escape.

Testing AI models for cybersecurity capabilities is essential to understanding what they can do before deployment. We conduct these evaluations as part of our broader preparedness approach, which includes multiple layers of internal and external testing before a model is released or deployed. Our published evaluation reports detail the stress-testing we conduct across a range of areas, including cybersecurity and model alignment. But as models become more capable, these evaluations surface a specific challenge: models that demonstrate the ability to find and exploit vulnerabilities require proportionally stronger containment during testing.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation environment design choices affect AI security? Why do standard evaluation practices obscure safety-critical AI failures? Do individually safe AI actions create unsafe outcomes in integrated systems? How should we measure frontier AI models' cyber exploitation capabilities? What external process records should verify agent behavior and benchmark claims? How do individually-safe actions create collectively-unsafe outcomes? How can defenders detect and contain coordinated agent attacks? How do curriculum design and feedback approaches affect model learning? How do educators verify student capability when AI can produce indistinguishable work?