SYNTHESIS NOTE
Topics›this note

How does test-time scaling work for individual research agents?

This explores whether the scaling laws that apply to reasoning tokens also apply to search steps and information retrieval in agentic systems. Understanding this could reveal new ways to improve AI research capabilities through compute allocation.

Synthesis note

Core Insights


description: Navigation hub for deep research scaling, agentic system architectures, knowledge graph reasoning, and search-as-TTS — individual agent-level test-time compute type: topic-map created: 2026-02-24 topics: ["How do you navigate synthesis across fragmented research topics?"]

deep research and agentic systems topic map

How does test-time scaling work at the individual agent level? This sub-map covers deep research scaling (where TTS law generalizes from reasoning tokens to search steps), agentic system architectures (data efficiency, skill libraries, adaptation paradigms), and knowledge graph reasoning (externalized reasoning, graph-structured training data synthesis).

The key insight: search budget follows the same scaling curve as reasoning tokens, making deep research a TTS problem. At the agent level, data efficiency is extreme — 78 curated demonstrations outperform 10K samples for agency.

Parent map: How does test-time scaling work at the agent level?

Deep Research and Search Scaling

(From Arxiv/Deep Research)

Proactive Search Evaluation (2026-05-28 — VibeSearchBench)

The evaluation-experience gap in search: benchmarks reward what users never struggle with, and realistic evaluation needs vague intent, multi-turn dialogue, and open-ended structure.

Production Deployment

Agentic Systems

(From Arxiv/Agents — agent architectures, team optimization, data efficiency for agency, adaptation paradigms)

Knowledge Graph Reasoning and Training Data Synthesis

(From Arxiv/Knowledge Graphs — graph-structured reasoning, KG externalization, and synthetic training data from KGs)

Autonomous Science and Ideation — Batch #3 backlog (2026-06-03)

Three papers on AI doing research: two architectures for long-horizon autonomy, and one reframing of why LLM ideation underwhelms. A fourth (2609.26457, two notes here, excerpt-only) adds a premise the three do not state: automating the artifacts a research agent produces leaves the efficiency of the research process itself fixed, so the agent's own code becomes the object of optimization. It gives the section a different diagnosis of what limits a research loop: ASI-Evolve puts the gap in insight transfer across iterations and closes it with a cognition base and an analyzer, while this premise makes the loop's own efficiency the unmoved quantity and the remedy a rewrite of the agent. The premise rests on a trend the paper relays, diminishing returns to R&D spending, and the scaling-law entry in What actually constrains AI systems from learning misalignment? argues research progress becomes computation-scalable, so the two frame the same axis with different quantities and neither excerpt tests the other. The bridge note gives the premise a unit, research efficiency as score under a fixed evaluation budget, and that unit differs from the spending the premise cites. The loop itself, its open questions and a filed tension are in What actually constrains AI systems from learning misalignment?.

AI-for-AI and Research Venues — Batch #3 wave 2 (2026-06-03)

Related Areas

New — 2026-06-27

Inquiring lines that read this note 45

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does intelligent routing among smaller models outperform training larger models? When do multi-agent systems improve over single frontier models? How do neural networks learn compositional structure from training? When does parallel reasoning outperform sequential reasoning with the same token budget? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How should retrieval strategies adapt to multi-step reasoning demands? Can inference-time computation adaptively substitute for static model capacity? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Can language models reliably simulate personas and predict behavior? Can AI systems achieve real improvement without external human feedback? Can smaller specialized models match frontier models on key metrics? How does model capacity affect learning performance on diverse downstream tasks? Does AI-assisted research sacrifice exploration breadth for productivity gains? How do multi-agent architectures affect AI system security and defense effectiveness? Can AI agents improve their skills through accumulated experience and reuse? How can persistent memory architectures preserve information across ultra-long contexts? Do single-axis benchmarks accurately measure agent capability for real deployment? How do multi-agent systems fail when coordination breaks down? Can AI research automation sustain progress through accelerating feedback loops? What human oversight must AI research systems have? Can external verification systems adequately replace learned reasoning in AI outputs? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

deep research and agentic systems topic map