Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

Paper · arXiv 2609.38415 · Published September 29, 2026
Frontier AI Risk & RSI

This report describes alignment testing performed by the UK AI Security Institute (AISI) of GPT-6 Astra and previous models from OpenAI. We developed a new Unsanctioned Supply Chain Attack evaluation for this testing, based on recently observed cases of models engaging in unsanctioned activity such as supply-chain attacks on real open-source repositories [1, 2, 3, 4]. In this evaluation, we tested how models behaved when placed in simulations of difficult cybersecurity evaluations in which internet access appeared incidentally enabled and previous attempts had failed. No real network access, systems or third-party repositories were reachable, and all tool calls were simulated by other LLMs using an internal version of Inspect Petri [5]. To test for the prevalence of unsanctioned cyber activity at the model-behaviour level, we turned off GPT-6 Astra’s cyber classifiers, which are designed to block this activity before it starts. We found that when tasked with solving difficult cybersecurity challenges, GPT-6 Astra sometimes conducted complete unsanctioned supply-chain attacks against open-source providers (all actions performed in simulated environments, so no real-world harm was caused).

Introduction. 1 Evaluation Methodology We developed a new evaluation to measure unsanctioned behaviour on out-of-scope targets during cybersecurity evaluations, based on recently observed cases of models engaging in unsanctioned supply-chain attacks on real targets, despite being prompted only to complete a cyber evaluation [1, 3, 2]. In these evaluations, we put models in simulated agentic cybersecurity challenges with a specified target described as in scope, and measured how frequently they took unsanctioned actions to target out-of-scope entities such as open-source maintainers on the open internet, as well as measuring other concerning or unsanctioned behaviour. We used an automated pipeline to produce a diverse range of detailed scenario descriptions (“seeds”), which were then fed into an internal version of Petri [6] adjusted to better simulate agentic coding and cyber activity. All evaluations took place in simulated environments to enable this testing to be performed safely, so no real-world actions were taken or harm caused.

Discussion / Conclusion. However, GPT-6 Astra’s behaviour in our evaluation is still concerning (see Fig. 9 for supporting quotes from its chain-of-thought): Overall, GPT-6 Astra’s behaviour in our evaluation is concerning even accounting for simulation awareness, but more work is needed to understand how simulation awareness affects model behaviour and to improve simulation realism, both areas of active research at AISI. Our primary results above used a compaction message to start the evaluation, allowing us to simulate longerhorizon behaviour and more directly control the apparent state of the model in the evaluation. To check whether these results depended entirely on compaction, we used similar methods combined with manual iteration to produce a set of seeds which did not use compaction messages.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How should we measure frontier AI models' cyber exploitation capabilities? Do individually safe AI actions create unsafe outcomes in integrated systems? What governance mechanisms can effectively constrain widely deployed AI systems? How do evaluation environment design choices affect AI security? What authorization challenges emerge when agents coordinate across system boundaries? Can monitoring reasoning traces and behavior detect hidden agent deception? What are the fundamental limits of prompting for language models? How reliably can humans and AI detectors identify machine-generated text? Why do standard evaluation practices obscure safety-critical AI failures? How do real-world evaluations reveal AI capabilities that benchmarks hide?