SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Can modular cognitive tools unlock reasoning without training?

Can reasoning capabilities be elicited by structuring LLM calls as isolated cognitive operations—understanding, recalling, examining, and backtracking—rather than through reinforcement learning?

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Cognitive architectures in psychology posit that reasoning arises from the orchestrated, sequential execution of modular, predetermined cognitive operations. The Cognitive Tools paper instantiates this in a modern tool-calling framework: four cognitive tools are implemented as discrete functions, each executed by the same LLM in a sandboxed context.

The four cognitive tools:

  1. Understand question: Breaks down the problem by identifying main concepts, extracting relevant information, highlighting properties/theorems/techniques that might help
  2. Recall related: Retrieves related knowledge of similar questions the model knows how to answer — guides reasoning through analogous examples
  3. Examine answer: Self-evaluation of a generated answer
  4. Backtracking: Returns to a prior reasoning state when a path appears unproductive

Unlike standard agentic tools (external APIs, calculators), cognitive tools encapsulate reasoning operations within the LLM itself. Each tool's schema includes a prompt template that isolates a specific cognitive operation; the LLM executes it in sandboxed context and feeds the structured result back into the main reasoning loop.

Results: GPT-4.1 on AIME2024 improves from 26.7% to 43.3% pass@1 — approaching o1-preview performance without any RL training. Similar gains across closed and open-weight models.

The key insight: modularity reduces interference between operations. Cognitive prompting (monolithic structured prompts) improves reasoning but lacks the isolation that makes modular cognitive architectures powerful. A tool-calling implementation enforces the sandboxed execution that pure prompting cannot guarantee.

This provides direct evidence for Do base models already contain hidden reasoning ability? — cognitive tools elicit pre-existing latent capability through structured invocation, not through training. The tool-calling framework is the elicitation mechanism.

The connection to Can structured argument prompts make LLM reasoning more rigorous?: both use structured decomposition of reasoning requirements to improve performance. Cognitive tools generalize this from argumentation-specific structure to domain-general cognitive operations.

Self-Discover as predecessor: Self-Discover (Zhou et al., 2024) is the clearest precursor to cognitive tools. It implements a two-stage process: (1) SELECT relevant atomic reasoning modules from a predefined set (critical thinking, step-by-step thinking, decomposition, etc.), (2) ADAPT selected modules to the specific task, (3) IMPLEMENT as a structured reasoning plan. The key difference from cognitive tools: Self-Discover composes a task-specific plan at inference time with only 3 extra inference steps — cheaper than the tool-calling loop but less modular. Self-Discover is more efficient (no sandboxed execution overhead) while cognitive tools provide stronger isolation between operations.

Inquiring lines that read this note 133

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What prevents LLMs from applying their reasoning knowledge to improve outputs? Why does polished AI output gain credibility despite fundamental verifiability problems? Why do language models struggle to implement user intent accurately from prompts? What causes coordination failures in multi-agent language model systems? Can language models reason beyond surface pattern matching? How does decomposing tasks into separate stages affect reasoning quality and safety? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can minimal training unlock latent reasoning already present in base models? Can latent reasoning match or exceed explicit reasoning performance? Can AI systems discover fundamental improvements to their own architectures? When do multi-agent systems improve over single frontier models? What prevents language models from performing systematic logical reasoning? Can inference-time computation adaptively substitute for static model capacity? Do accumulated memories help or hurt continual learning in models? What are the fundamental limits of prompting for language models? Can LLMs distinguish between linguistic form and semantic meaning? Does augmenting symbolic reasoning improve LLM logical reasoning ability? How reliably can language models perform causal versus temporal reasoning? What limits language model accuracy in evaluating ideas? Can language models reliably simulate personas and predict behavior? Can mechanistic interpretability methods reliably reveal what models actually know? How do curriculum design and feedback approaches affect model learning? Should models ask for clarification when facing ambiguous or under-specified information? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How can we reduce inherent biases in LLM-based evaluation judges? Does training data format shape model reasoning more than domain content? Can AI systems achieve real improvement without external human feedback? What prediction granularity best trains models to generate reliable reasoning? How do individually-safe actions create collectively-unsafe outcomes? How do thinking tokens exhibit diminishing returns in reasoning? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? How do reward signal properties affect model reasoning and safety? Does pretraining establish the ceiling for what reward learning can improve? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How do philosophical assumptions about AI consciousness affect practical harms and design? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Should agents compress episodic memory or retain raw interaction histories? Can reasoning traces reveal actual model reasoning versus plausible output? How much of agent capability comes from harness versus the model itself? Why don't better reasoning capabilities improve theory of mind performance? Does AI assistance help or harm professional skill development?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 205 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cognitive tools implement reasoning operations as modular agentic tool calls that elicit reasoning without rl training