INQUIRING LINE

When an AI writes out its step-by-step reasoning, is that a real record of its thinking, or a convincing story?

How do we measure genuine reasoning inside a language model?

This explores how researchers can tell whether a language model is actually reasoning or just producing text that looks like reasoning, and which kinds of tests and measurements tell the two apart.


This explores how researchers can tell whether a language model is actually reasoning or just producing text that looks like reasoning. The corpus suggests a surprising starting point: the step-by-step reasoning a model writes out is a poor place to look. When researchers feed models reasoning traces with invalid logical steps, or train them on corrupted traces, performance barely changes. That suggests the written trace is closer to a persuasive performance than a record of what the model computed Do reasoning traces show how models actually think?. Chain-of-thought seems to work by steering the model into familiar reasoning patterns from its training data. You can see this when the problems shift away from that data: performance drops in a predictable way, which is what imitation looks like, not a new capability Does chain-of-thought reasoning reveal genuine inference or pattern matching?.

If the written reasoning can't be trusted, one option is to look inside the model. A measure called the deep-thinking ratio counts how many tokens have their predicted value change substantially as they pass through the model's layers. That ratio tracks accuracy on hard math and science benchmarks. It is also useful in practice: picking answers by this signal matches the accuracy of sampling many answers and taking a vote, at lower cost Can we measure how deeply a model actually reasons?. Looking inside also cuts the other way. Models trained to output meaningless filler tokens in place of a written chain of thought still compute the correct answer in their early layers, then overwrite it before producing output. The reasoning can be real but hidden, and only layer-by-layer analysis finds it Do transformers hide reasoning before producing filler tokens?.

A second approach is to design tests that separate reasoning from things that resemble it. Keep the logical rules but remove the meaningful content (swap real-world concepts for arbitrary symbols), and performance collapses. This shows that models lean on learned associations rather than formal logic Do large language models reason symbolically or semantically?. Vary how unfamiliar the specific problem is rather than how hard it is, and failures follow unfamiliarity: models fit patterns from instances they've seen rather than learning general procedures Do language models fail at reasoning due to complexity or novelty?. The same split between surface success and real ability appears in social reasoning. Models pass structured tests of understanding what others believe but fall back on surface strategies in open-ended conversation Do large language models genuinely simulate mental states?. Adding irrelevant padding to inputs also cuts reasoning accuracy from 92% to 68% at only about 3,000 tokens, far below the model's context limit Does reasoning ability actually degrade with longer inputs?.

Measurement can also mislead you in the other direction, by making a model look worse at reasoning than it is. Some well-known "reasoning collapses" turn out to be execution failures. The model knows the algorithm but can't carry out hundreds of steps in plain text, and when given tools it solves problems past the supposed limit Are reasoning model collapses really failures of reasoning?. Producing reasoning and judging it are also separate skills. Frontier models that solve problems almost perfectly score as low as 48% when asked to spot flawed steps in solutions that reach the right answer Can models that reason well also grade reasoning well?. So a model that gets the right answer isn't necessarily one that can tell good reasoning from bad.

The practical takeaway: there is no single test for "genuine reasoning." What the corpus offers is a set of cross-checks. Look inside the model rather than at its written explanation. Change the meaningful content and the familiarity of problems while keeping the logic fixed. And separate reasoning from execution limits and from the ability to verify reasoning. One related signal: a model's confidence in its own answers can serve as a training reward that improves step-by-step reasoning and also keeps the model's confidence better matched to its accuracy, without human labels Can model confidence work as a reward signal for reasoning?. That hints the model holds some internal signal about the quality of its reasoning even when its written explanation doesn't show it.


Sources 11 notes

Do reasoning traces show how models actually think?

LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.

Does chain-of-thought reasoning reveal genuine inference or pattern matching?

CoT works by constraining models to reproduce familiar reasoning patterns from training, not by enabling novel symbolic reasoning. Performance degrades predictably under distribution shifts—the signature of imitation rather than capability emergence.

Can we measure how deeply a model actually reasons?

Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.

Do transformers hide reasoning before producing filler tokens?

Logit lens analysis shows models trained with hidden CoT tokens compute correct answers in layers 1-3, then actively suppress these representations in final layers to produce format-compliant filler output. The reasoning is fully recoverable from lower-ranked token predictions.

Do large language models reason symbolically or semantically?

When semantic content is decoupled from reasoning tasks, LLM performance collapses even with correct rules in context. Models rely on parametric commonsense and token associations rather than formal logical manipulation, constraining reasoning to training distribution semantics.

Show all 11 sources
Do language models fail at reasoning due to complexity or novelty?

LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.

Do large language models genuinely simulate mental states?

ChangeMyView and FANTOM benchmarks show LLMs fail at authentic perspective-taking in open-ended scenarios, despite succeeding on structured tasks. Hybrid Bayesian architectures that force explicit belief tracking outperform LLM-alone approaches, suggesting the gap is architectural rather than merely training-based.

Does reasoning ability actually degrade with longer inputs?

FLenQA shows reasoning accuracy drops from 92% to 68% at just 3000 tokens of padding, far below context window capacity. The degradation is task-agnostic, uncorrelated with language modeling performance, and persists even with chain-of-thought prompting.

Are reasoning model collapses really failures of reasoning?

Models confined to text-only generation cannot execute multi-step procedures at scale, even when they know the underlying algorithm. Tool-enabled models solve problems beyond the supposed reasoning cliff, suggesting the bottleneck is procedural execution bandwidth.

Can models that reason well also grade reasoning well?

Frontier reasoning models solve problems near-perfectly but score as low as 48% when grading solutions with correct answers but flawed steps. Outcome-focused training rewards answer production, not step-by-step verification, leaving evaluation starved.

Can model confidence work as a reward signal for reasoning?

RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.