INQUIRING LINE

Inside a language model, does recalling facts happen in different layers than reasoning, and can sharpening reasoning blur what it knows?

Do different network layers specialize in retrieval versus reasoning tasks?

This explores whether, inside a language model, the work of pulling up stored facts happens in different layers than the work of reasoning over them, and what follows if it does.


This explores whether, inside a language model, recalling stored facts and reasoning over them happen in different layers. The corpus answers yes in one direct finding. In Why does reasoning training help math but hurt medical tasks?, inference is modeled in two phases: lower layers retrieve knowledge and higher layers adjust the reasoning. This split has a practical consequence. Reasoning training mostly reshapes the upper layers, which helps on math, where the facts are few and the procedure matters most. The same training can hurt in medicine, where answers depend on recalling a large body of specific facts accurately. So a model that 'reasons better' can know less in practice, because the tuning that sharpened its procedure disturbed how it reaches its knowledge.

The layer split matches a split in where the model learned each skill. Does procedural knowledge drive reasoning more than factual retrieval? traced model behavior back to 5 million pretraining documents. Reasoning drew on broad, transferable 'how-to' knowledge spread across many sources. Factual recall depended on narrow memorization of particular documents. Read together, the two notes suggest that retrieval and reasoning differ in how they are learned as well as where they sit in the network. Facts are stored item by item, while procedures are general patterns that carry over to new problems. This may explain why Can non-reasoning models catch up with more compute? finds that training, not extra compute at answer time, separates reasoning models from the rest. The procedure has to be trained into the model. Giving it more time to think does not create it.

There is a caution. Specialized layers do not mean the reasoning layers run a clean, general algorithm. Do language models fail at reasoning due to complexity or novelty? and Does chain-of-thought reasoning actually generalize beyond training data? both find that reasoning breaks down on unfamiliar instances and on data unlike what the model was trained on. It does not break down simply because a problem is harder. The 'reasoning' half of the split may still lean on recognizing patterns the model has seen before.

System designers outside the model have arrived at a similar separation. Do hierarchical retrieval architectures outperform flat ones on complex queries? finds that splitting query planning from answer writing reduces interference on questions that need several retrieval steps. Can routing queries to task-matched structures improve RAG reasoning? goes further and routes each query to the kind of knowledge structure that suits it, such as a table, a graph, or plain text chunks. There is also evidence that mixing the two jobs causes harm. Can relevant memories actually harm LLM reasoning? shows that accurate, relevant memories can still lower reasoning performance below a no-memory baseline. Retrieved material can disrupt reasoning even when it is correct.

The corpus has only one note that studies layers directly, so treat the clean lower-versus-higher picture as a well-supported model rather than settled anatomy. What the collection adds is the pattern around it. The same separation shows up in how the model learns (memorized facts versus general procedures), in how its layers are organized, and in how engineers build retrieval systems. In each place, mixing the two jobs causes problems that keeping them apart avoids.


Sources 8 notes

Why does reasoning training help math but hurt medical tasks?

Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.

Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

Can non-reasoning models catch up with more compute?

Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.

Do language models fail at reasoning due to complexity or novelty?

LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.

Does chain-of-thought reasoning actually generalize beyond training data?

DataAlchemy experiments show CoT fails systematically under distributional shifts in task, length, and format. Models produce fluent but logically inconsistent reasoning — imitating reasoning form without valid underlying logic.

Show all 8 sources
Do hierarchical retrieval architectures outperform flat ones on complex queries?

Separating query planning from answer synthesis into distinct components reduces interference and improves multi-hop query performance. This architectural principle mirrors documented benefits of separating planning from execution in agent design.

Can routing queries to task-matched structures improve RAG reasoning?

StructRAG demonstrates that selecting knowledge structure type based on query demands—via DPO-trained router choosing among tables, graphs, algorithms, catalogues, and chunks—improves knowledge-intensive reasoning over standard retrieval. The approach grounds this in cognitive load and cognitive fit theory from cognitive science.

Can relevant memories actually harm LLM reasoning?

MemTrapBench shows that all five tested memory frameworks underperform a no-memory baseline, with drops exceeding 10%, despite memories being accurately stored and task-relevant. This reveals a failure mode at the point of use that standard memory benchmarks miss.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.