SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can algorithms control LLM reasoning better than LLMs alone?

Explores whether embedding LLMs within algorithmic control flow—where programs manage state and context filtering—enables complex task decomposition beyond what LLMs achieve through self-managed reasoning chains.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

LLM Programs embed an LLM within an algorithm rather than asking the LLM to be the algorithm. The critical design choice: instead of the LLM maintaining the current state of the program (its context), the LLM is presented with only step-specific prompt and context for each step. A classic computer program (Python) handles control flow, parsing of outputs, and augmentation of prompts for succeeding steps.

This is distinct from both Chain-of-Thought (where the LLM manages state through its token stream) and agentic frameworks (where the LLM decides what to do next). In LLM Programs, the algorithm structure is external and explicit, not learned or generated:

The key benefit is information hiding. By concealing information irrelevant to the current step, each LLM call focuses on an isolated subproblem whose results feed future calls. This addresses two fundamental limitations:

  1. Capability limits: Complex tasks that are currently too difficult because they require coordinating multiple reasoning steps
  2. Architectural constraints: The finite context window restricts processing to what fits within it

The approach recognizes the LLM as a limited general agent and avoids further training. Instead, the expected behavior is recursively deconstructed into simpler steps the LLM can perform to a sufficient degree.

This connects to Can modular cognitive tools unlock reasoning without training? — both decompose reasoning into modular operations. But LLM Programs are more structured: the control flow is predetermined by the algorithm, whereas cognitive tools are flexibly invoked. It also extends Does separating planning from execution improve reasoning accuracy? — the program IS the decomposer, and each LLM call IS the solver, with clean separation enforced by architecture rather than training.

Decomposed Prompting as the software library formalization: Decomposed Prompting (Khot et al., 2022) makes the software library analogy explicit. The decomposer defines a top-level program using interfaces to simpler sub-task functions. Sub-task handlers serve as "modular, debuggable, and upgradable implementations" — if a particular handler underperforms, it can be debugged in isolation, replaced with an alternative prompt or even a symbolic system (e.g., Elasticsearch), and plugged back in. This is more general than least-to-most prompting: it supports recursive decomposition, non-linear structures, and mixed neural-symbolic pipelines. The key architectural insight is that sub-task handlers are shared across tasks, creating a reusable prompt library — the closest existing analog to how software engineers build with functions. Source: Prompts Prompting.

Inquiring lines that read this note 169

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can language models reliably simulate personas and predict behavior? Can LLMs distinguish between linguistic form and semantic meaning? Does intelligent routing among smaller models outperform training larger models? What prevents LLMs from applying their reasoning knowledge to improve outputs? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How should agents coordinate through shared persistent code artifacts? What makes process supervision effective for training complex reasoning models? How do curriculum design and feedback approaches affect model learning? Why does AI verification capability persistently exceed generation capability? Can AI systems discover fundamental improvements to their own architectures? What human oversight must AI research systems have? Does augmenting symbolic reasoning improve LLM logical reasoning ability? What prevents language models from performing systematic logical reasoning? Should GUI agents use structured screen representations instead of end-to-end vision? How does decomposing tasks into separate stages affect reasoning quality and safety? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Do language models reason through disagreement or only accommodate it? What causes coordination failures in multi-agent language model systems? How should humans and AI agents share control and decision-making? Why do language models fail at sustained therapeutic relationships despite understanding techniques? Why do retrieval-augmented generation systems fail in practice despite sound architecture? How effectively can test-time voting aggregate diverse reasoning samples? Can minimal training unlock latent reasoning already present in base models? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Can inference-time computation adaptively substitute for static model capacity? What makes agent memory systems durable and reusable across sessions? Why do planning and grounding require opposing optimization strategies? How do individually-safe actions create collectively-unsafe outcomes? Why do language models struggle to implement user intent accurately from prompts? Can language models reason beyond surface pattern matching? Can mechanistic interpretability methods reliably reveal what models actually know? When does parallel reasoning outperform sequential reasoning with the same token budget? What are the fundamental limits of prompting for language models? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How does diversity prevent model convergence on superficial patterns? When do multi-agent systems improve over single frontier models? How reliably can language models perform causal versus temporal reasoning? Can reasoning traces reveal actual model reasoning versus plausible output? How can persistent memory architectures preserve information across ultra-long contexts? How should retrieval strategies adapt to multi-step reasoning demands? Can AI systems evade safety evaluations through reasoning manipulation? Should governance of agentic AI systems be runtime or design-time? How do neural networks learn compositional structure from training? Why do models reveal hidden associations despite concealment attempts? What explains the gap between benchmark scores and true reasoning capability? Should agents compress episodic memory or retain raw interaction histories? Can smaller specialized models match frontier models on key metrics? How do multi-agent systems fail when coordination breaks down? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Why do language models hallucinate and how can we prevent it? Why do standard evaluation practices obscure safety-critical AI failures? How should systems validate code that agents generate? What prediction granularity best trains models to generate reliable reasoning? Can AI agents improve their skills through accumulated experience and reuse? How much of agent capability comes from harness versus the model itself? Can external verification systems adequately replace learned reasoning in AI outputs? Do individually safe AI actions create unsafe outcomes in integrated systems? How can we reduce inherent biases in LLM-based evaluation judges? What authorization challenges emerge when agents coordinate across system boundaries? How can models maximize welfare while preserving minority veto rights? Does AI assistance help or harm professional skill development? What gaps exist between benchmark performance and real deployment outcomes? How do clinicians calibrate trust in AI medical recommendations? What limits language model accuracy in evaluating ideas? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 136 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM programs decompose complex tasks into step-specific prompts within algorithmic control flow — hiding irrelevant context per step