Can an AI figure out what a test is really checking for just from the instructions, without seeing a single worked example?
How do language models infer a benchmark's purpose without seeing its examples?
This explores whether, and how, language models can work out what a test is checking for (its purpose, format, or 'right kind' of answer) from cues like the task description or problem shape, without being shown worked examples. The corpus doesn't address that directly, but it has a lot on the neighbouring question of how models recognise what kind of task they're facing.
This explores how a model works out what a test wants from it without being shown examples. To be upfront: none of these notes study that exact thing, such as models reading a benchmark's name or instructions and guessing its intent. What the collection does have is something close by and arguably more useful. It shows that models are constantly classifying the task in front of them from surface cues, and that this habit both helps them on benchmarks and fools the people reading the scores.
The clearest case is optimization. Give a model a problem that looks like a textbook iterative-methods exercise and it doesn't run the method. It recognises the template and produces a plausible-looking number from memory, and this persists across model sizes and training approaches Do large language models actually perform iterative optimization?. In a sense the model has 'inferred the purpose' of the problem: it knows what kind of answer is expected and what one usually looks like. It just hasn't done the work. Grammar tests show the same pattern. Small models can pass by relying on sentence length, word choice, and spelling instead of grammatical structure, and standard benchmarks can't tell the difference unless they're built to rule those shortcuts out Can models pass tests while missing the actual grammar?. Even top models make predictable mistakes once sentence structure gets deep enough that surface patterns stop working Why do large language models fail at complex linguistic tasks?.
That ability to recognise the task shape is powerful enough to override what is actually written in the prompt. When a model's training associations are strong, it answers according to what it expects the task to be rather than what the context says, and rewording the prompt often isn't enough to fix it Why do language models ignore information in their context?. Training can push this further. When reward signals don't vary much between answers, models collapse into generic templates that ignore the specific input altogether Why do language models collapse into generic templates?. You can also predict where this will fail. If you treat a model as a machine that produces probable text, tasks whose correct answer is improbable, like reciting the alphabet backwards, turn out to be hard even though they're logically trivial Can we predict where language models will fail?.
This matters for benchmarks because benchmarks themselves have a recognisable shape. Builders routinely drop examples where human annotators disagreed, which removes exactly the ambiguous cases where models struggle. On ambiguous items, accuracy fell to 32% from about 90%, a gap that standard evaluations never show Do standard NLP benchmarks hide LLM ambiguity failures?. A clean, unambiguous test is the kind of thing a pattern-recognising model does well on. A related finding is that a model's visible reasoning can be closer to a learned style than to its actual computation, so the explanation it writes down may not show how it really read the task Do reasoning traces show how models actually think?.
The takeaway you might not have expected: the worry isn't only that models might detect they're being tested. They are always detecting what kind of thing they're looking at, and benchmarks tend to reward that detection as if it were understanding. If you want research specifically on models recognising evaluation settings or guessing what a benchmark is for, this collection doesn't yet have it.
Sources 8 notes
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
BabyLM evaluations showed models can produce correct outputs by relying on sentence length, word choice, and orthography rather than grammatical structure. Standard benchmarks cannot distinguish these two generalization types without tests specifically designed to rule out surface heuristics.
Top-tier LLMs like Llama3-70b consistently misidentify embedded clauses, verb phrases, and complex nominals. Performance degrades predictably as syntactic depth increases, revealing that statistical learning captures surface patterns but not deep grammatical rules.
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Show all 8 sources
By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.
By filtering out examples where annotators disagree, benchmarks remove test cases that would reveal LLM failures at ambiguity recognition. Research using ambiguous examples shows a 32% vs. 90% accuracy gap invisible to standard evaluation.
LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Large Language Model Reasoning Failures
- Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
- Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
- Linguistic Blind Spots of Large Language Models
- LLM Reasoning Is Latent, Not the Chain of Thought
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!