SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Can reasoning systems forget history without losing coherence?

Does treating each reasoning step as independent—rather than accumulating historical context—actually preserve problem-solving quality while reducing computational waste? This explores whether Markov-style memoryless reasoning can scale effectively.

Synthesis note · 2026-05-18 · sourced from Reinforcement Learning

Existing test-time scaling methods all carry history along. Chain-based methods preserve the entire reasoning trace to generate each next step. Tree-based methods track ancestor and sibling relationships across branches. Graph-based methods compound this with arbitrary node dependencies. As reasoning scales, the accumulated historical dependencies waste compute and — worse — interfere with the model's ability to reason effectively on the current state.

Atom of Thoughts (2502.12018) makes a different bet: each reasoning state should be a simplified problem equivalent to the original, with partial reasoning steps either transformed into known conditions or excluded as incorrect explorations. The state transition mechanism has two phases. First, decompose the current question into a dependency-based directed acyclic graph (DAG) capturing structural information. Second, contract the subquestions into a new independent question. Iterate the decomposition-contraction until reaching directly-solvable atomic questions.

The Markov property is the load-bearing claim. Each transition depends only on the current state — never on the path that produced it. This is not a heuristic; it is a structural property guaranteed by answer-equivalence preservation through contraction. If the contracted question yields the same answer as the original, no historical context is required to continue.

The cognitive science motivation is direct. Humans solve complex problems by identifying and resolving self-evident subquestions, then reformulating a simplified problem state — not by maintaining detailed reasoning processes for resolved components. The reformulation IS the memory management.

Two architectural advantages emerge. AoT eliminates the need for maintaining and computing historical information when scaling test-time compute, and atomic questions can be seamlessly integrated into existing TTS frameworks as a plug-in enhancement. Since Can recursive subtask trees overcome context window limits?, AoT is the language-level version of the same insight — TIMRUN prunes KV cache to free positional embeddings; AoT contracts subproblems to free conceptual context. Both reject the assumption that more history equals better reasoning.

Inquiring lines that read this note 114

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should designers communicate what AI systems truly are and can do? Does AI assistance promote real skill development or substitute for independent learning? What happens to knowledge when intelligence becomes tokenized like a commodity? Can models improve accuracy without degrading reasoning quality? What causes reasoning models to fail or wander off track? How do prompting refinements mask underlying biases and model frequency patterns? Can compression size predict model complexity better than parameter count alone? What design and behavioral factors drive false consciousness attribution to AI? How should systems decide whether to retrieve or reason alone? Can multi-agent systems avoid converging on false agreement without deliberation? How should inference compute be allocated based on problem difficulty? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How does reasoning length affect model performance across different tasks? What reasoning architectures enable models to solve complex problems efficiently? What attack surfaces do reasoning traces and chains introduce? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Why does adding new knowledge through fine-tuning degrade existing capabilities? Why don't LLMs reliably translate capability into accurate outputs? Do reasoning traces faithfully reflect actual model reasoning? Why does memory consolidation cause performance regression in continual learning? How should agents manage memory granularity to improve long-term performance? Can inference-time compute effectively substitute for model scale? Can intelligent routing over smaller models outperform scaling a single large model? What capability trade-offs arise from domain specialization through fine-tuning? How effectively can language models perform reasoning, especially combined with symbolic methods? What structural properties of attention create systematic model biases? Can brute-force automated research substitute for iterative depth and human research intuition? How should test-time compute scaling work in agentic systems? Can memory architectures handle ultra-long context better than attention? Should agents decouple planning from perception grounding for better performance? Can diffusion models match autoregressive performance on language generation tasks? How can evolutionary algorithms maintain diversity during solution search? How do surface patterns enable correct outputs but reduce robustness? How should retrieval systems handle complex multi-step reasoning? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? How does decomposing tasks improve reasoning and prevent failure propagation? Can self-generated feedback reliably guide model training without ground truth? Is reasoning capability latent in base models or created by post-training? Do reasoning benchmarks predict model performance in long-horizon workflows? How do standardized protocols improve multi-agent coordination and reliability? How do prompt design choices influence model reasoning and performance?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 133 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

markov-style memoryless reasoning replaces accumulated-history test-time scaling with iterative decompose-then-contract