SYNTHESIS NOTE
Topics›Linguistics, NLP, NLU›this note

Are models actually reasoning about constraints or just defaulting conservatively?

Do language models genuinely apply constraints when solving problems, or do they simply prefer harder options by default? Minimal pair testing reveals whether apparent reasoning success masks hidden biases.

Synthesis note · 2026-05-01 · sourced from Linguistics, NLP, NLU

The Heuristic Override Benchmark uses minimal pairs — same surface heuristic, with versus without the implicit constraint — to test whether apparent reasoning successes reflect actual reasoning. The result is striking. Twelve of fourteen models perform worse on the no-constraint variant than on the constraint-active variant, with drops up to 38.5 percentage points. Only two models (GPT-OSS-120B at +13.8 and GPT-OSS-20B at +11.0) improve when the constraint is removed.

This exposes a hidden mechanism behind apparent accuracy. When the constraint is present, the correct answer is the harder one (drive to the car wash that is 50m away). When the constraint is removed, the correct answer is the easier one (walk to the store that is 50m away). Models that default to recommending the harder option score correctly on constraint-active cases without doing any constraint reasoning. They are not solving the problem. They are reflexively choosing the more conservative option, which happens to coincide with the constraint-required answer.

The minimal-pair asymmetry is the only test that catches this. Single-instance accuracy looks fine — the model recommended driving, the right answer was driving. But the same model recommends driving even when walking would be correct, because the recommendation is not based on the constraint. The two-of-fourteen models that improve on minimal pairs are the only ones whose constraint-active accuracy reflects genuine reasoning about the constraint. The rest are riding a conservative-bias accident that aggregate metrics cannot distinguish from reasoning.

Inquiring lines that read this note 131

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can smaller specialized models match frontier models on key metrics? What makes reasoning traces effective supervision even when they're incorrect? What explains the gap between benchmark scores and true reasoning capability? Why do training associations persist despite contradictory contextual information? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can inference-time computation adaptively substitute for static model capacity? What prevents language models from performing systematic logical reasoning? How does diversity prevent model convergence on superficial patterns? How does RLHF training shape models to prioritize agreement over accuracy? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What are the fundamental limits of prompting for language models? Why does polished AI output gain credibility despite fundamental verifiability problems? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Can reasoning models use reflection to correct their initial outputs? Should models ask for clarification when facing ambiguous or under-specified information? What limits language model accuracy in evaluating ideas? Can language models reason beyond surface pattern matching? Can reasoning traces reveal actual model reasoning versus plausible output? How susceptible are language models to conversational persuasion and belief change? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What prevents LLMs from applying their reasoning knowledge to improve outputs? Do persona-based approaches introduce systematic biases in user simulation? Can base models hide emergent misalignment through alignment training? How do reward models systematically fail to represent diverse human preferences? How do thinking tokens exhibit diminishing returns in reasoning? How reliably can language models perform causal versus temporal reasoning? How does model capacity affect learning performance on diverse downstream tasks? Why does self-revision amplify confidence in wrong model answers? Can mechanistic interpretability methods reliably reveal what models actually know? What distinguishes genuine communicative competence from surface language performance? How do reward signal properties affect model reasoning and safety? What prediction granularity best trains models to generate reliable reasoning? How can evaluations be made robust against model reward hacking? How does fine-tuning trade off accuracy against reasoning quality? How do individually-safe actions create collectively-unsafe outcomes? Should governance of agentic AI systems be runtime or design-time? Can we trust AI-generated mathematical proofs without understanding them? Can external verification systems adequately replace learned reasoning in AI outputs? What gaps exist between benchmark performance and real deployment outcomes? How does awareness of evaluation context influence model behavior?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Conservative bias hides behind apparent reasoning success — most models perform worse when the constraint is removed than when it is present