When an AI explains its reasoning step by step, is that the real story — or just a plausible-sounding one made up after the fact?
How faithfully do model scratchpads reflect the actual reasons behind model decisions?
This explores whether the step-by-step reasoning a model writes out before answering (its scratchpad or chain of thought) is a true record of how it reached its decision, or a story told alongside that decision.
This explores whether a model's written-out reasoning is a real log of how it decided, or a plausible account produced next to the decision. The corpus mostly points to the second answer. When researchers compared the logical structure a trace presents with the model's internal cause-and-effect pathways, the two didn't match. Most of the wrong steps in a trace had no effect on the final answer Do reasoning traces actually show how models think?. The trace looks like the path to the answer, but the answer often doesn't depend on it.
The strongest evidence comes from experiments that break the reasoning on purpose. Models trained on deliberately corrupted or irrelevant traces solve problems about as well as models trained on correct ones, and sometimes they generalize better to unfamiliar problems Do reasoning traces need to be semantically correct?. Logically invalid steps perform nearly as well as valid ones Do reasoning traces show how models actually think?. For DeepSeek's R1, the 'thinking' tokens are generated the same way as any other output and have no special status as computation. Invalid traces often still lead to correct answers, which suggests the trace is closer to learned style than to working reasoning Do reasoning traces actually cause correct answers?. A related finding fits this picture: a small model given only lightweight fine-tuning matched much larger models on reasoning tasks. That suggests much of what 'reasoning training' teaches is how to format and organize output Can small models reason well by just learning output format?.
This matters most for safety. If people plan to oversee models by reading their scratchpads, that approach can fail in two ways Can we actually trust reasoning model outputs?. The first is omission: whatever actually drove the decision never shows up in the trace. The second is laundering: questionable reasoning appears in the trace, but in clean, harmless-sounding language. A similar pattern shows up outside reasoning. When editing documents, stronger models introduce subtle errors while keeping the text looking intact, where weaker models visibly delete content Does model capability change how documents degrade?. More capable models can produce output that looks fine and still isn't.
The twist is that traces still help, just not by being accurate explanations. Training on messy exploration, including dead ends and backtracking, makes reasoning more robust Can models learn better by training on messy exploration paths?. Models can also improve by keeping only the self-generated rationales that reached the right answer, without checking whether the steps were valid Can models improve by filtering only on answer correctness?. So the extra tokens seem to act as working space for computation, not as a faithful description of it.
If the text doesn't show the real reasoning, where do you look? One answer is to look inside the model. The 'deep-thinking ratio' measures how often the model's prediction for a token changes substantially as it passes through the model's layers. That internal signal tracks accuracy better than the length or appearance of the trace Can we measure how deeply a model actually reasons?. The takeaway: a scratchpad is better treated as a tool the model uses than as a confession. To learn what actually drove a decision, measure what happens inside the model, or check the output against something external such as tests or proofs When can weak models match strong model performance?.
Sources 11 notes
ReasoningFlow found that most erroneous steps in traces don't influence final answers, and critically, the discourse structure traces present linguistically does not match their actual internal causal pathways. This gap suggests traces are narrative surface rather than verified computation logs.
Models trained on systematically irrelevant traces maintain solution accuracy and sometimes improve out-of-distribution generalization, suggesting traces function as computational scaffolding rather than meaningful reasoning steps.
LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.
R1's intermediate tokens carry no special execution semantics and are generated identically to other LLM output. Invalid traces frequently produce correct answers, proving traces are not causally necessary—they correlate with answers via learned formatting, not functional reasoning.
A 1.5B parameter model with LoRA-only post-training matched larger full-parameter RL models on reasoning tasks, suggesting RL teaches output format organization rather than new factual knowledge. This efficiency indicates reasoning and knowledge storage are separable capabilities.
Show all 11 sources
Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Research shows that training on messy trajectories—failed attempts, self-correction, and backtracking—teaches more robust reasoning than training only on shortcut solutions. This approach models o1-style deep reasoning as search internalization rather than solution memorization.
STaR demonstrates that self-generated rationales filtered exclusively by answer correctness improve reasoning performance significantly. On CommonsenseQA, this correctness-filtered approach achieved 72.5% accuracy, outperforming direct answer fine-tuning and closing the gap with models 30 times larger.
Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.
Sampling alone amplifies coverage but cannot select correct solutions. Reliable performance matching requires external soundness signals—tests, proofs, or type checks—that convert latent correct proposals into actual selections.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Do Reasoning Representations Help Humans Evaluate LLM Outputs?
- Large Language Model Reasoning Failures
- LLM Reasoning Is Latent, Not the Chain of Thought
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think