Does the reasoning cliff depend on how we test models?
If language models hit a capability wall in text-only reasoning tasks, does that wall disappear when they can use tools? What does this reveal about what we're actually measuring?
Apple's "Illusion of Thinking" identifies three regimes of reasoning model performance: (1) easy tasks solved reliably, (2) a narrow zone of genuine reasoning improvement, and (3) catastrophic failure beyond a complexity threshold — the reasoning cliff. This finding generated significant attention as evidence that LLM reasoning is fundamentally limited.
The agentic reframe: When the same models are evaluated with tool access (code execution, search, verification), the cliff disappears. Performance continues scaling beyond the text-only collapse point. The "reasoning cliff" is actually a tool-absence cliff — a composite measurement of reasoning ability and execution capability, where execution becomes the bottleneck at higher complexity.
Why this matters: Text-only evaluation creates a specific lens that conflates two separable abilities. A model may correctly identify the reasoning strategy but fail to execute it in pure text (tracking multiple variables, maintaining state, performing sequential calculations). Tool access offloads execution, revealing the reasoning capability that was always present.
The evaluation implication: Benchmarks that prohibit tool use measure something real but not what they claim. They measure text-only reasoning+execution, not reasoning capability. For deployment decisions — where models will typically have tool access — text-only evaluations systematically underestimate capability.
This connects to Why do reasoning LLMs fail at deeper problem solving? — which may be partly an execution failure mode rather than a reasoning failure mode. It also connects to Are reasoning model collapses really failures of reasoning?: reasoning models that seem to fail at hard problems may actually fail at hard execution while succeeding at hard reasoning.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can smaller specialized models match frontier models on key metrics? What prevents language models from performing systematic logical reasoning? What explains the gap between benchmark scores and true reasoning capability? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Can minimal training unlock latent reasoning already present in base models? How does model capacity affect learning performance on diverse downstream tasks? How does fine-tuning trade off accuracy against reasoning quality?Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Can Large Language Models Reason and Optimize Under Constraints?
- Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
- LLM Reasoning Is Latent, Not the Chain of Thought
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Original note title
the reasoning cliff is evaluation-boundary-dependent — text-only assessment shows capability collapse that disappears in agentic tool-enabled settings