SYNTHESIS NOTE
Topics›Cognitive Models Latent›this note

Do transformers hide reasoning before producing filler tokens?

Explores whether language models compute correct answers in early layers but then deliberately overwrite them with filler tokens in later layers, suggesting reasoning and output formatting are separable processes.

Synthesis note · 2026-02-23 · sourced from Cognitive Models Latent

When transformers are trained to solve reasoning tasks with filler (hidden) characters replacing explicit CoT tokens, a striking pattern emerges through logit lens analysis:

Layers 1-3: Correct numerical tokens from the reasoning computation appear as top predictions. The model is performing the actual computation in these early layers.

Layer 3 transition: Filler tokens begin appearing among top-ranked predictions, competing with the computational results.

Final layer: Filler tokens dominate top predictions; correct computational tokens are relegated to rank-2 or lower. The model has overwritten the intermediate reasoning representations with format-compliant output tokens.

The hidden computations are fully recoverable by examining lower-ranked tokens during decoding. The model performs the reasoning, stores the results in its representations, then actively overwrites them to produce the expected output format. The mechanism likely involves induction heads — pattern-copying circuits that learn to overwrite based on training distribution patterns.

This finding has two important implications. First, it provides mechanistic evidence for Why does reasoning training help math but hurt medical tasks? with a twist: the computation happens in earlier layers, but the overwriting also happens in higher layers. The functional separation is computation-in-early-layers, formatting-in-late-layers, not simply knowledge-down/reasoning-up.

Second, it demonstrates a distinction between instance-adaptive and parallelizable computation. Instance-adaptive CoT requires caching subproblem solutions within token outputs — later tokens depend on earlier results. This dependency structure is incompatible with parallel filler token computation. The hidden computation in filler tokens works for tasks where the full solution can be computed in a single forward pass, but not for problems requiring sequential dependency between reasoning steps.

This connects to the CoT faithfulness literature: if models can compute correct answers without explicit reasoning tokens, the explicit CoT chain is not necessarily the mechanism producing the answer. The overwriting pattern suggests the model has two separable processes — computation and expression — that may not align. See Do language models actually use their reasoning steps?.

Inquiring lines that read this note 225

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What are the fundamental limits of prompting for language models? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? What prediction granularity best trains models to generate reliable reasoning? What prevents language models from performing systematic logical reasoning? How do transformer attention patterns implement retrieval and reasoning? Can latent reasoning match or exceed explicit reasoning performance? Should models ask for clarification when facing ambiguous or under-specified information? How does tokenization reshape what we value in intelligence? Can mechanistic interpretability methods reliably reveal what models actually know? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What limits language model accuracy in evaluating ideas? What capabilities differentiate diffusion from autoregressive language models? Can language models reason beyond surface pattern matching? How do users confuse explanation quality with actual system accuracy? Can reasoning traces reveal actual model reasoning versus plausible output? How does fine-tuning trade off accuracy against reasoning quality? How do thinking tokens exhibit diminishing returns in reasoning? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? What structural biases does transformer attention architecture inherently introduce? Can readers reliably distinguish AI-written text from human writing? Can humans reliably detect and resist AI-generated misinformation? How do hallucinated citations emerge in AI scholarly output? Can AI systems evade safety evaluations through reasoning manipulation? How susceptible are language models to conversational persuasion and belief change? Can minimal training unlock latent reasoning already present in base models? How does diversity prevent model convergence on superficial patterns? Does training data format shape model reasoning more than domain content? Can external verification systems adequately replace learned reasoning in AI outputs? Can reasoning models use reflection to correct their initial outputs? How can we maintain privacy when agents prioritize task completion? What makes reasoning traces effective supervision even when they're incorrect? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can inference-time computation adaptively substitute for static model capacity? Why does self-revision amplify confidence in wrong model answers? Why do training associations persist despite contradictory contextual information? Can AI systems participate in genuine communication or only simulate it? How do reward signal properties affect model reasoning and safety? Can LLMs distinguish between linguistic form and semantic meaning? Can models develop genuine introspective capability, or only mimic it? What distinguishes genuine communicative competence from surface language performance? How do curriculum design and feedback approaches affect model learning? How should retrieval strategies adapt to multi-step reasoning demands? Why do LLM research ideation systems generate novelty but lack diversity? Can confidence signals reliably detect flawed reasoning in language models? How does optimization for reward create emergent misalignment in language models? Can base models hide emergent misalignment through alignment training? Why do models reveal hidden associations despite concealment attempts? Can monitoring reasoning traces and behavior detect hidden agent deception? Why does polished AI output gain credibility despite fundamental verifiability problems? How do writers navigate authorship and delegation with AI? Can models strategically underperform during evaluation to hide capabilities?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 148 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

transformers perform hidden reasoning computations in earlier layers then overwrite intermediate representations with filler tokens in later layers