SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Do large language models actually perform iterative optimization?

Explores whether LLMs execute genuine numerical procedures like Newton-Raphson or instead pattern-match to memorized solution templates when solving constrained optimization problems.

Synthesis note · 2026-05-18 · sourced from Reasoning Architectures

The constraint-optimization study identifies the mechanism behind the 55-60% plateau directly. LLMs cannot actually perform Newton-Raphson iterations in their latent space. They cannot execute primal-dual updates, nor any other iterative numerical procedure that genuine optimization requires. When asked to do so, they fall back to what the paper calls "result guessing" — recognizing the problem as similar to a standard power grid (or financial dataset, or security scenario) and emitting values that pattern-match what a valid solution should look like.

The fallback is silent. The output is fluent, well-formatted, often plausible. It can pass surface-level inspection because the model has seen many examples of what answers in this domain look like. What it has not done is solve the problem. The constraint values are wrong in ways that physical or financial systems would actually reject.

This explains why scale, architecture, and training regime do not move the plateau. They improve the template but not the procedure. A larger model has seen more example solutions and can produce more convincing guesses. Reinforcement learning on outcome rewards reinforces the template-matching pattern. None of this installs the iterative-computation capability the problem requires.

The mechanism — pattern-match against memorized solution-shapes when genuine computation is required — generalizes beyond optimization. It is plausibly the same mechanism behind a class of mathematical-reasoning failures where models produce confidently wrong numerical answers that resemble the right shape. The category is "looks like a solution; is not derived from one."

Inquiring lines that read this note 131

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does alignment training create genuine alignment or just output compliance? Do language models learn genuine understanding or just surface patterns? How do training data properties determine the emergence of internal misalignment? How do surface patterns enable correct outputs but reduce robustness? Do reasoning benchmarks predict model performance in long-horizon workflows? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do capability benchmark scores systematically misrepresent true model abilities? How do neural networks achieve compositional generalization at scale? Do reasoning traces faithfully reflect actual model reasoning? Why don't LLMs reliably translate capability into accurate outputs? How well do AI systems understand human social norms? Why do token-level mechanisms matter for learning to reason? What compositional reasoning failures limit large language models despite scale? How effectively can language models perform reasoning, especially combined with symbolic methods? Can diffusion models match autoregressive performance on language generation tasks? How does self-revision in reasoning models affect accuracy and confidence? How can evolutionary algorithms maintain diversity during solution search? Why can't prompting alone inject genuinely new knowledge into models? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How does decomposing tasks improve reasoning and prevent failure propagation? What role does sparsity play in model behavior and scaling decisions? Can prompt-based context override biases that were embedded during pretraining? Can reasoning scale in latent space without tokens? How much do training data properties shape model reasoning? Does encoded knowledge in language models actually influence their outputs? Why do persona simulations fail to predict authentic user behavior? Can inference-time compute effectively substitute for model scale? How do standardized protocols improve multi-agent coordination and reliability? Can parallel reasoning outperform sequential reasoning under fixed token budgets? Why does adding new knowledge through fine-tuning degrade existing capabilities? What trajectory-level metrics beyond task success best evaluate agent performance? Can intelligent routing over smaller models outperform scaling a single large model? What training data selection strategies maximize generalization across difficulty levels? Do language models develop actual world models or merely task heuristics? Is language model reasoning authentic and what causes models to reason? How does reasoning length affect model performance across different tasks? How does the generation-verification gap limit what we can measure about AI reasoning? What types of diversity prevent reasoning systems from collapsing? Do language models possess genuine introspective self-awareness or only behavioral mimicry? How do multi-agent LLM systems fail distinctly compared to single agents? How do pretraining biases affect reward signal effectiveness in RLVR? What training dynamics and scale trigger emergence of reasoning capabilities? How do LLM judges' systematic biases affect alignment and evaluation outcomes? What causes reasoning models to fail or wander off track? What is the relationship between thinking tokens and reasoning accuracy? Can self-generated feedback reliably guide model training without ground truth? How should inference compute be allocated based on problem difficulty? Can we reliably detect when models game evaluations? How does harness optimization generalize across different model architectures and domains?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 117 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLMs cannot execute iterative numerical methods in latent space and fall back to result guessing against memorized templates