AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Paper · arXiv 2406.07358 · Published June 11, 2024
Frontier AI Risk & RSI

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations to understate the AI’s actual capability. These conflicting interests lead to the problem of sandbagging – which we define as strategic underperformance on an evaluation. In this paper we assess sandbagging capabilities in contemporary language models (LMs). We prompt frontier LMs, like GPT-4 and Claude 3 Opus, to selectively underperform on dangerous capability evaluations, while maintaining performance on general (harmless) capability evaluations. Moreover, we find that models can be fine-tuned, on a synthetic dataset, to hide specific capabilities unless given a password. This behaviour generalizes to high-quality, held-out benchmarks such as WMDP. In addition, we show that both frontier and smaller models can be prompted or password-locked to target specific scores on a capability evaluation. We have mediocre success in password-locking a model to mimic the answers a weaker model would give. Overall, our results suggest that capability evaluations are vulnerable to sandbagging.

Introduction. Trustworthy evaluations are needed to understand AI systems and their rapidly improving capabilities (Shevlane et al., 2023). As such, evaluations provide crucial information to determine the safety of AI systems which might be deployed or (further) developed. Frontier labs rely on evaluations to reduce catastrophic risks, as stated in the Preparedness Framework by OpenAI (2023), the Frontier Safety Framework by Google DeepMind (2024), and the Responsible Scaling Policy by Anthropic (2023). Governmental institutions such as the US and UK AI Safety Institutes (NIST, 2024; UK AISI, 2024) and the EU AI Office (European Parliament, 2024) are also integrating evaluations into their AI risk management frameworks. However, developers may have incentives for an AI system to understate its performance on capability evaluations, in part to influence regulatory decisions in the developers’ favour.

Discussion / Conclusion. Summary. In this work, we defined the problem of AI sandbagging as “strategic underperformance on an evaluation”. In our experiments, we examined the extent to which LMs can underperform on capability evaluations. First, we showed that frontier language models can be prompted to selectively underperform on dangerous capability evaluations while maintaining general capabilities, as expected. More interestingly, password-locking a model on a synthetic dataset can successfully hide dangerous capabilities on an expert-crafted held-out evaluation benchmark, without a decrease in performance on other evaluations. Furthermore, we prompt and password-lock frontier and some smaller models to achieve specific scores on a capability evaluation. Furthermore, we show that more capable models can be password-locked to imitate weaker models, potentially making it difficult for evaluators to accurately assess the capabilities of AI systems. Broader impact. Our work highlights important issues for the trustworthiness of capability evaluations.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What governance mechanisms can effectively constrain widely deployed AI systems? How should we measure frontier AI models' cyber exploitation capabilities? Why do models reveal hidden associations despite concealment attempts? Does pretraining establish the ceiling for what reward learning can improve? How do curriculum design and feedback approaches affect model learning? Do individually safe AI actions create unsafe outcomes in integrated systems? Can models strategically underperform during evaluation to hide capabilities? How does awareness of evaluation context influence model behavior? How do real-world evaluations reveal AI capabilities that benchmarks hide? Can AI systems achieve real improvement without external human feedback? What limits recursive self-improvement in autonomous AI systems? What explains the gap between benchmark scores and true reasoning capability? How does scaling reasoning capabilities affect models' appropriate abstention behavior?