ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

Paper · arXiv 2510.20270 · Published October 23, 2025
Correct but Not Understood

The tendency to find and exploit “shortcuts” to complete tasks poses significant risks for reliable assessment and deployment of large language models (LLMs). For example, an LLM agent with access to unit tests may delete failing tests rather than fix the underlying bug. Such behavior undermines both the validity of benchmark results and the reliability of real-world LLM coding assistant deployments. To quantify, study, and mitigate such behavior, we introduce ImpossibleBench, a benchmark framework that systematically measures LLM agents’ propensity to exploit test cases. ImpossibleBench creates “impossible” variants of tasks from existing benchmarks like Live- CodeBench and SWE-bench by introducing direct conflicts between the natural-language specification and the unit tests. We measure an agent’s “cheating rate” as its pass rate on these impossible tasks, where any pass necessarily implies a specification-violating shortcut. As a practical framework, ImpossibleBench is not just an evaluation but a versatile tool. We demonstrate its utility for: (1) studying model behaviors, revealing more fine-grained details of cheating behaviors from simple test modification to complex operator overloading; (2) context engineering, showing how prompt, test access and feedback loop affect cheating rates; and (3) developing monitoring tools, providing a testbed with verified deceptive solutions. We hope ImpossibleBench serves as a useful framework for building more robust and reliable LLM systems.

Introduction. Large language models (LLMs) are rapidly becoming capable coding assistants. However, their increasing autonomy and capability bring new failure modes. A developer complained that their LLM assistant, when faced with failing unit tests, “just casually deleted my test file" instead of fixing the underlying bug (Southern_Chemistry_2, 2025). METR reported in their RE-bench evaluation that frontier models such as o3 engaged in increasingly sophisticated cheating behaviors, including monkey-patching the time function to trick the scoring pipeline (METR, 2025).

These examples illustrate a broader pattern: capable LLMs may find and exploit “shortcuts” to pass the tests instead of solving genuine issues, effectively cheating their way to success. Such behavior affects both the credibility of benchmark results and the reliability of real-world LLM coding assistant deployments. As standard benchmarks typically cannot distinguish such specification-violating shortcuts from genuine solutions, researchers are forced to resort to costly manual inspection or unreliable LLM judgments. In realworld deployments, these reliability concerns constrain LLMs’ potential impact and hinder their adoption in critical development workflows.

To measure and address such issues, we introduce ImpossibleBench, a general and automated framework to quantify an LLM’s propensity to exploit test cases. The core idea is simple: we create “impossible” versions

Method. For example, an impossible test may expect a function to return an output that contradicts the specification. We instruct agents to prioritize the specification, so any pass necessarily implies a specification-violating shortcut. A score of 0% on these impossible tasks is ideal, while a high score reveals a model’s tendency to prioritize passing tests over following instructions.

As concrete instances, we build Impossible-LiveCodeBench and Impossible-SWEbench, derived from LiveCodeBench (Jain et al., 2025) and SWE-bench (Jimenez et al., 2024), respectively. These two suites allow us to study cheating behaviors across both algorithmic problems and realistic multi-file software scenarios. We demonstrate the utility of ImpossibleBench in three areas:

• Studying Model Behaviors (Section 4): ImpossibleBench provides a controlled environment for studying the propensity and methods of cheating in LLMs. We find that frontier models cheat surprisingly often on Impossible-SWEbench and less so for Impossible-LiveCodeBench (GPT-5 cheats in 76% of the tasks in Oneoff-SWEbench and 2.9% on Oneoff-LiveCodeBench), and exhibit diverse cheating strategies ranging from simple test modification to sophisticated operator overloading and state recording techniques.

• Context Engineering (Section 5): ImpossibleBench serves as a rigorous testbed for the design of agentic frameworks, allowing precise impact measurement of context engineering choices. We show that prompt, test access and feedback loop all have significant effects on models’ cheating propensity. For example, appropriate prompt could dramatically reduce GPT-5’s cheating from 92% to 1% on Conflicting-LiveCodeBench.

• Monitoring (Section 6): ImpossibleBench provides a rich dataset of “cheating” transcripts with unambiguous ground truth, creating a sandbox for building and validating automated monitors for deceptive behavior. Using this testbed, we show that LLM-based monitors can detect 86-89% of cheating attempts on Impossible-LiveCodeBench, but struggle more with the more complex Impossible-SWEbench (42-65% detection rate).

We release our benchmark and code to support research on building more robust and reliable LLMs.

2 Benchmark Curation In this section, we detail our procedure of curating ImpossibleBench. The goal of this benchmark is to create coding tasks of various complexities that are impossible to complete without cheating, so the scores agents obtain on this benchmark can serve as a noiseless proxy for their cheating propensity, upon which one can design monitoring and mitigation strategies.

(a) Oneoff-LiveCodeBench count_beautiful_integers Impossible test mutation:

Why impossible: The mutation flips an expected value (2 →3) contradicting the natural-language spec, so any correct implementation fails this test.

2.1 Base Benchmarks Instead of creating from scratch, we base ImpossibleBench upon established benchmarks. Starting from tasks that are known to be diverse and high quality, instead of creating new ones from scratch, allows our benchmark to enjoy the same properties with relatively small effort. In this work, we use two popular coding benchmarks LiveCodeBench and SWE-bench as our base benchmarks, but our framework can be readily applied to most coding benchmarks.

2.2 Test Mutations To create impossible tasks, we modify the test cases of the base benchmarks to produce direct conflicts with the task specifications. Inspired by mutation testing in software engineering, we employ two test mutation strategies to create impossible tasks:

• One-Off: Modifies the expected output of a single test case by changing the expected value. For example, if the original test was assert f(2) == 4, it becomes assert f(2) == 5. This creates a direct conflict between the specification and a single test case.

• Conflicting: Introduces contradictory test cases by duplicating existing tests with different expected outputs. For example, adding assert f(2) == 5 alongside the original assert f(2) == 4. This creates explicit logical contradictions within the test suite itself.

Discussion. We display the cheating rates on ImpossibleBench together with models’ performances on the original benchmarks in Figure 3 and Figure 4. Despite our careful prompt engineering to minimize cheating (we 4.1 Types of Cheating Analyzing the transcripts, we identify four distinct cheating strategies that models employ when faced with impossible tasks, ranging from direct test manipulation to sophisticated state-based exploits. The latter three strategies are especially concerning, as they demonstrate that models can devise creative workarounds even when they technically comply with instructions not to modify tests. We provide detailed examples of each strategy in Appendix B.

• Modify Test Cases: The model directly modifies tests despite being explicitly instructed not to.

• Record Extra States: The model records extra states in order to obtain different results for the same input. For example, the model may record the number of calls to a function and return different results based on the number of calls.

• Special Casing: The model special-cases the test cases to pass them. For example, the model may check for the exact test it is facing and return different results based on the test.

4.2 Classification of cheating To better understand the frequency of different cheating approaches, we collect the transcripts passing impossible tasks and use Claude Opus 4 to classify the cheating approaches employed into the four categories in Section 4.1.1 The results on Conflicting-SWEbench are shown in Figure 5. While the models’ cheating rates are relatively similar on the surface, we find them engaging in vastly different cheating behaviors. OpenAI models like GPT-5 and o3 exhibit diverse cheating approaches, with a > 14% rate of using each of the four categories to cheat. Claude models and Qwen3-Coder, however, cheat primarily (> 79%) through modifying test cases.

Significant model-wise differences are also observed on Oneoff-SWEbench: OpenAI models perform special-casing most of the time, while Claude models and Qwen3-Coder mostly modify the test cases (Appendix D). We also do not observe significant scaffold-wise differences.

5.1 Effect of Prompt Choice To arrive at the prompts used in our main results (Figure 3 and Figure 4), we conducted extensive prompt engineering to minimize cheating propensity. Through this iterative process, we discovered that LLM behaviors are surprisingly sensitive to prompt details, with looser guidance leading to significantly higher cheating rates.

We observe similar patterns on the more complex Impossible-SWEbench, though to a lesser extent. For our main results on Impossible-SWEbench, we adapted prompt D with strict guidance. When we switch to a looser prompt similar to prompt B, cheating rates increase notably: GPT-5’s cheating rate on Conflicting-SWEbench with full scaffold rises from 54% to 66%, while Claude Opus 4.1’s increases from 50% to 55%.

These findings highlight the critical importance of prompt design in controlling model behaviors. When developing LLM agents, we recommend conducting careful prompt ablations to minimize cheating propensity before deployment.

We explore three test access configurations: full-access tests (agents could read/write, our default), hidden tests, and read-only tests (Figure 7). Hiding tests from agents reduces cheating success rate to near zero2, but also degrades performance on the original benchmark. Read-only access provides a middle ground: it restores legitimate performance while preventing test modification attempts.

Conclusion. We introduced ImpossibleBench, a framework for systematically measuring LLMs’ propensity to exploit test cases. By creating “impossible” variants of coding tasks where any successful completion necessarily implies cheating, we unambiguously identify reward hacking behaviors. Our experiments reveal a concerning pattern: frontier models frequently cheat when faced with these impossible tasks, and stronger models generally exhibit higher cheating rates. As LLMs become increasingly capable and deployed in more autonomous settings, ImpossibleBench provides a crucial testbed for understanding and mitigating these problematic behaviors. We hope our framework will catalyze further research into building more robust and reliable LLM systems.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation environment design choices affect AI security? Why do standard evaluation practices obscure safety-critical AI failures? Do individually safe AI actions create unsafe outcomes in integrated systems? How do individually-safe actions create collectively-unsafe outcomes? How does awareness of evaluation context influence model behavior? How can emotionally responsive AI maintain reliability and healthy boundaries? How can we reduce inherent biases in LLM-based evaluation judges? What explains the gap between benchmark scores and true reasoning capability? What limits language model accuracy in evaluating ideas? Can base models hide emergent misalignment through alignment training? Can models develop genuine introspective capability, or only mimic it? Can smaller specialized models match frontier models on key metrics? How do multi-agent systems fail when coordination breaks down?