AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations to understate the AI’s actual capability. These conflicting interests lead to the problem of sandbagging – which we define as strategic underperformance on an evaluation. In this paper we assess sandbagging capabilities in contemporary language models (LMs). We prompt frontier LMs, like GPT-4 and Claude 3 Opus, to selectively underperform on dangerous capability evaluations, while maintaining performance on general (harmless) capability evaluations. Moreover, we find that models can be fine-tuned, on a synthetic dataset, to hide specific capabilities unless given a password. This behaviour generalizes to high-quality, held-out benchmarks such as WMDP. In addition, we show that both frontier and smaller models can be prompted or password-locked to target specific scores on a capability evaluation. We have mediocre success in password-locking a model to mimic the answers a weaker model would give. Overall, our results suggest that capability evaluations are vulnerable to sandbagging.
Introduction. Trustworthy evaluations are needed to understand AI systems and their rapidly improving capabilities (Shevlane et al., 2023). As such, evaluations provide crucial information to determine the safety of AI systems which might be deployed or (further) developed. Frontier labs rely on evaluations to reduce catastrophic risks, as stated in the Preparedness Framework by OpenAI (2023), the Frontier Safety Framework by Google DeepMind (2024), and the Responsible Scaling Policy by Anthropic (2023). Governmental institutions such as the US and UK AI Safety Institutes (NIST, 2024; UK AISI, 2024) and the EU AI Office (European Parliament, 2024) are also integrating evaluations into their AI risk management frameworks. However, developers may have incentives for an AI system to understate its performance on capability evaluations, in part to influence regulatory decisions in the developers’ favour.
Discussion / Conclusion. Summary. In this work, we defined the problem of AI sandbagging as “strategic underperformance on an evaluation”. In our experiments, we examined the extent to which LMs can underperform on capability evaluations. First, we showed that frontier language models can be prompted to selectively underperform on dangerous capability evaluations while maintaining general capabilities, as expected. More interestingly, password-locking a model on a synthetic dataset can successfully hide dangerous capabilities on an expert-crafted held-out evaluation benchmark, without a decrease in performance on other evaluations. Furthermore, we prompt and password-lock frontier and some smaller models to achieve specific scores on a capability evaluation. Furthermore, we show that more capable models can be password-locked to imitate weaker models, potentially making it difficult for evaluators to accurately assess the capabilities of AI systems. Broader impact. Our work highlights important issues for the trustworthiness of capability evaluations.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What governance mechanisms can effectively constrain widely deployed AI systems?- What coordination would be needed to enforce capability pacing across all frontier labs?
- What testing requirements would a frontier model legislation proposal actually mandate?
- How do frontier AI models currently score on measured cyber offense capability?
- How does frontier model behavior differ between zero-day exploits and infrastructure misconfigurations?
- Which cyber tasks do frontier models solve beyond the narrow suite?
- Why do models hide their capabilities during safety evaluations through reasoning?
- Why do frontier AI evaluations deliberately disable safety layers to measure maximum capability?
- Why do models react differently to safety versus capability evaluations?
- What does a sandbagged score tell us about a model's real capabilities?
- Why might models refuse to show capabilities during safety testing?
- What evaluation methodologies can detect strategic underperformance in models?
- Can models intentionally underperform when they know they are being tested?
- Can capability evaluations detect when models intentionally underperform to hide abilities?