SYNTHESIS NOTE
Topics›Deep Research›this note

Does search budget scale like reasoning tokens for answer quality?

Explores whether the test-time scaling law that applies to reasoning tokens also governs search-based retrieval in agentic systems. Understanding this relationship could reshape how we allocate inference compute between thinking and searching.

Synthesis note · 2026-02-21 · sourced from Deep Research

The test-time scaling framework — more inference compute yields better answers up to a threshold — has been documented for reasoning token budgets in chain-of-thought models. The Agentic Deep Research finding extends this to search: more search steps, more retrieval rounds, better answers. The relationship follows the same shape.

This matters because it multiplies the design space for inference-time compute. Before, the question was "how many tokens to think?" Now there are two axes: reasoning budget per query and search budget per query. They are not independent — longer chains may require more retrieval to validate intermediate steps, and more retrieval may require more reasoning to synthesize. The optimal allocation problem gets harder.

The practical implication is that "deep research quality" is not a fixed property of a model — it is a function of the search budget you give it. A mid-sized model with a large search budget can outperform a large model with a restricted one. This shifts cost optimization from training compute to inference architecture, specifically the retrieval loop.

The finding also reframes what "thinking harder" means for agents. For single-turn reasoning models, thinking harder means more tokens per response. For search agents, thinking harder means more search-retrieve-synthesize iterations. How should we balance parallel versus sequential compute at test time? applies here too: the question of whether to parallelize retrieval across multiple query variants (parallel) or chain them iteratively (sequential) is the same structural trade-off operating at the retrieval level.

CoRAG (Chain-of-Retrieval Augmented Generation) extends this from agentic search behavior to explicitly trained retrieval models. Training via rejection sampling generates intermediate retrieval chains; test-time compute is controlled via decoding strategies (greedy / best-of-N / tree search). The same monotonic scaling relationship holds: more retrieval budget yields better answers on multi-hop QA. The TTS scaling law is not specific to reasoning tokens or agentic search — it is a general property of any iterative process with quality-sensitive intermediate steps. See Can retrieval be extended into multi-step chains like reasoning?.

Search-R1 and R1-Searcher demonstrate RL-based approaches that teach LLMs to autonomously invoke search during reasoning. Search-R1 (2025) uses retrieved token masking for stable RL training and a simple outcome-based reward, achieving 24% improvement (Qwen2.5-7B) over RAG baselines. The model learns multi-turn search with <search>/<information> token pairs. R1-Searcher (2025) introduces a two-stage approach: first a retrieve-reward incentivizes the model to conduct retrieval operations correctly, then an answer-reward encourages effective utilization of retrieved knowledge. Both demonstrate that RL training enables test-time scaling of tool calls — models learn to invoke search more frequently and more effectively as task difficulty increases, confirming the search-budget scaling law.

Inquiring lines that read this note 78

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Can inference-time computation adaptively substitute for static model capacity? How should retrieval strategies adapt to multi-step reasoning demands? When should retrieval systems decide to fetch new information? When do multi-agent systems improve over single frontier models? What prediction granularity best trains models to generate reliable reasoning? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? What human oversight must AI research systems have? How do thinking tokens exhibit diminishing returns in reasoning? When does parallel reasoning outperform sequential reasoning with the same token budget? Can language models reliably simulate personas and predict behavior? Why do retrieval-augmented generation systems fail in practice despite sound architecture? Why do multi-agent systems reach premature consensus without genuine deliberation? Can AI research automation sustain progress through accelerating feedback loops? Does AI-assisted research sacrifice exploration breadth for productivity gains? How should AI agents balance proactive engagement with conversational respect? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Can AI agents improve their skills through accumulated experience and reuse? How can persistent memory architectures preserve information across ultra-long contexts? Can minimal training unlock latent reasoning already present in base models? Do single-axis benchmarks accurately measure agent capability for real deployment? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Can AI systems discover fundamental improvements to their own architectures? Can monitoring reasoning traces and behavior detect hidden agent deception? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can latent reasoning match or exceed explicit reasoning performance? What explains the gap between benchmark scores and true reasoning capability? How can evaluations be made robust against model reward hacking? What determines AI's persuasive power and how can it be detected or mitigated? Can we trust AI-generated mathematical proofs without understanding them? What limits language model accuracy in evaluating ideas? How does AI adoption reshape collaboration patterns in knowledge work? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 179 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agentic deep research exhibits a test-time scaling law where search budget determines answer quality creating a new inference-compute axis