Does giving an AI search and outside tools make it think better, or just let it reach more facts?
Do knowledge access methods like search improve reasoning or just coverage?
This explores whether giving a model access to outside knowledge (web search, retrieval, tools, knowledge graphs) makes it reason better, or only lets it reach more facts while its thinking stays the same.
This explores whether search and retrieval change how well a model thinks, or only how much it can reach. The corpus suggests the line between the two is blurrier than the question assumes. Some kinds of outside access only widen coverage. Others change which problems a model can solve at all. A good place to start is inside the model. One study finds that factual knowledge sits mostly in the lower layers of a network, while reasoning adjustments happen in the higher layers Why does reasoning training help math but hurt medical tasks?. That separation helps explain why reasoning training can help math and still hurt medicine: knowing and reasoning really are different jobs. A related finding from pretraining data points the same way. Reasoning draws on broad, reusable *how-to* knowledge from many sources, while recalling a fact depends on having memorized that specific fact Does procedural knowledge drive reasoning more than factual retrieval?. Plain fact lookup mostly fills in the second kind, which is coverage.
Tools make a stronger case. A formal proof shows that letting a model call tools such as code execution expands what it can reason about. Some strategies are impossible, or far too long-winded, to carry out in text alone, and the gain shows up on abstract reasoning, not only arithmetic Do tools actually expand what language models can reason about?. Another note argues that when reasoning models suddenly 'collapse' on harder puzzles, they often still know the algorithm. They just can't carry out hundreds of steps in text. Given tools, they solve problems past the supposed cliff Are reasoning model collapses really failures of reasoning?. So part of what looks like better reasoning is really execution capacity being handed off to the tool.
How the knowledge is shaped matters as much as how much of it there is. StructRAG sends each question to the kind of structure that suits it (a table, a graph, an algorithm, or plain text chunks) and beats one-size-fits-all retrieval on knowledge-heavy reasoning Can routing queries to task-matched structures improve RAG reasoning?. SymAgent derives rules from a knowledge graph's structure and uses them as navigation plans, which beats retrieval that matches text by similarity alone Can symbolic rules from knowledge graphs guide complex reasoning?. In both cases the retrieved material does more than supply facts: it gives the model a structure to think through.
Here is the surprising part. Search budget behaves like thinking budget. In agentic deep research, answer quality rises with more search rounds and then levels off, the same curve seen with more reasoning tokens. That makes search a second dial you can trade against thinking Does search budget scale like reasoning tokens for answer quality?. The two dials can also compete. If a model thinks without limit inside one search turn, it uses up the context it needs to absorb the next round of evidence, so per-turn reasoning limits help Does limiting reasoning per turn improve multi-turn search quality?.
One caution: access can't fix a model that searches badly in its own head. Reasoning models tend to wander. They explore without system and drop promising paths too early, and their success rate falls off exponentially as problems get deeper Why do reasoning LLMs fail at deeper problem solving? Why do reasoning models abandon promising solution paths?. Spending effort on a range of high-level approaches rather than ever-longer chains helps Can abstractions guide exploration better than depth alone?. The corpus doesn't directly test plain web-search RAG for reasoning gains against coverage gains. What it does suggest is that outside knowledge improves reasoning when it adds execution power or structure, and mostly adds coverage when it only adds facts.
Sources 11 notes
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.
Formal proof shows tool-integrated reasoning enables strategies impossible or prohibitively verbose in text alone, expanding both empirical and feasible support. The advantage spans abstract reasoning, not just arithmetic, and Advantage Shaping Policy Optimization stabilizes training without reward distortion.
Models confined to text-only generation cannot execute multi-step procedures at scale, even when they know the underlying algorithm. Tool-enabled models solve problems beyond the supposed reasoning cliff, suggesting the bottleneck is procedural execution bandwidth.
StructRAG demonstrates that selecting knowledge structure type based on query demands—via DPO-trained router choosing among tables, graphs, algorithms, catalogues, and chunks—improves knowledge-intensive reasoning over standard retrieval. The approach grounds this in cognitive load and cognitive fit theory from cognitive science.
Show all 11 sources
SymAgent derives symbolic rules from KG structure using LLM reasoning to create navigational plans that align natural language with graph topology. This approach captures structural reasoning patterns explicitly, outperforming retrieval methods that rely on semantic similarity alone.
Agentic deep research shows monotonic-to-diminishing-returns curves for search iterations, matching reasoning token scaling. This creates a new inference-compute axis: models can trade off reasoning budget against search budget to optimize answer quality.
Unrestricted reasoning within single search turns consumes context needed for subsequent retrieval rounds, degrading the agent's ability to incorporate new evidence. Setting per-turn reasoning budgets, not just overall time limits, prevents this context erosion and maintains search quality across iterations.
Current reasoning models lack the three properties of systematic exploration: validity, effectiveness, and necessity. This causes success probability to drop exponentially with problem depth, making medium problems solvable but deep problems catastrophically harder.
Reasoning LLMs exhibit two reinforcing failures: wandering (invalid exploration) and underthinking (premature path-switching). Decoding-level interventions like thought-switching penalties improve accuracy without fine-tuning, suggesting viable solutions exist but are abandoned prematurely.
RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- Reasoning LLMs are Wandering Solution Explorers
- Large Language Model Reasoning Failures
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Efficient Tool Use with Chain-of-Abstraction Reasoning