The Illusion of the Illusion of the Illusion of Thinking

Paper · Source
LLM Failure ModesReasoning Critiques

"Shojaee et al.’s underlying observations hint at a more subtle, yet real, challenge for LRMs: a brittleness in sustained, high-fidelity, step-by-step execution.

The true illusion is the belief that any single evaluation paradigm can definitively distinguish between reasoning, knowledge retrieval, and pattern execution.

"inadvertently shift the goalposts of what is being measured.... the truth is more nuanced than either “fundamental failure” or “simple artifact.” ...while the dramatic “collapse” is indeed an illusion, the original experiments, when viewed through the lens of the critique, still point toward genuine and important limitations in the execution of complex, sequential tasks."

solution length (Shojaee et al.’s primary metric for complexity) is not equivalent to computational difficulty.

  1. The Fragility of Sustained Execution: The fact that models fail at

high-iteration sequential tasks, even when the underlying logic is simple

(like Tower of Hanoi), points to a weakness in sustained, step-by-step

processing. While the hard token limit is the ultimate cause of failure in the

experiment, the enormous token cost itself is a symptom of how LLMs represent

and execute such problems. A system with more robust internal state-tracking

might execute the steps more efficiently.

  1. Unexplained Behavioral Ticks: Opus

and Lawsen’s critique does not fully account for one of Shojaee et al.’s most

intriguing findings: that near the collapse point, LRMs

“begin reducing their reasoning effort (measured by inference-time tokens)”. If

 the issue were merely hitting a hard output limit, one might expect models to

 consistently reason until that limit is reached. This counter-intuitive

 decline in effort on harder problems, also noted in other contexts

 [3], suggests a more complex behavioral scaling property that warrants further

 investigation.

  1. Data Contamination and Generalization: Shojaee et al.

 observed that models could handle a 100+ move Tower of Hanoi problem but

 failed a far shorter River Crossing problem. They speculate this is due to the

 prevalence of the former in training data. This highlights a key challenge in

 evaluation: distinguishing true, generalizable reasoning from sophisticated

 pattern matching of familiar problems, a core issue in compositionality

 [4]. Opus and Lawsen’s “generate a function” test, when applied to a very

 common problem, falls into this same ambiguity.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do single-axis benchmarks accurately measure agent capability for real deployment? What prevents LLMs from applying their reasoning knowledge to improve outputs? What makes reasoning traces effective supervision even when they're incorrect? Can mechanistic interpretability methods reliably reveal what models actually know? How does decomposing tasks into separate stages affect reasoning quality and safety? How do thinking tokens exhibit diminishing returns in reasoning? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What prevents language models from performing systematic logical reasoning? Can minimal training unlock latent reasoning already present in base models? Can reasoning traces reveal actual model reasoning versus plausible output? How does fine-tuning trade off accuracy against reasoning quality? How does diversity prevent model convergence on superficial patterns? Why does self-revision amplify confidence in wrong model answers? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Why do multi-agent systems reach premature consensus without genuine deliberation? How does model capacity affect learning performance on diverse downstream tasks? Can language models reason beyond surface pattern matching?