SYNTHESIS NOTE
Topics›Test Time Compute›this note

Do critique models improve diversity during training itself?

Explores whether critique integrated into the training loop, beyond test-time scoring, actively maintains solution diversity and prevents the model from converging too narrowly during iterative self-training.

Synthesis note · 2026-02-20 · sourced from Test Time Compute

The intuitive framing of critique models is that they help at test time: the model generates, the critic scores, we select the best. But the more important finding from AutoMathCritique is that critique integrated into the training loop improves the actor model's exploration efficiency and solution diversity during training itself.

Without critique in the loop, iterative self-training suffers from "tail narrowing" — the model converges on a narrow distribution of solutions, becoming less able to explore diverse reasoning paths. The critique model counteracts this: by providing step-level feedback on exploration, it guides the actor toward high-quality paths it wouldn't have discovered alone, maintaining distributional breadth through training.

This connects to Does policy entropy collapse limit reasoning performance in RL?: critique models are a way to maintain entropy — the exploration needed for continued improvement — without relying solely on architectural entropy management (Clip-Cov, KL-Cov). The critique is an external signal that prevents premature convergence.

The implication: critique models are training infrastructure as much as inference infrastructure. Evaluating them only on test-time accuracy misses their more fundamental role.

Inquiring lines that read this note 103

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? What are the fundamental limits of prompting for language models? Why do LLM research ideation systems generate novelty but lack diversity? Can minimal training unlock latent reasoning already present in base models? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Why do training associations persist despite contradictory contextual information? What limits recursive self-improvement in autonomous AI systems? When does parallel reasoning outperform sequential reasoning with the same token budget? Can readers reliably distinguish AI-written text from human writing? Can AI systems achieve real improvement without external human feedback? How do reward signal properties affect model reasoning and safety? How do training data quality and composition affect downstream model performance? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Can smaller specialized models match frontier models on key metrics? Can AI agents improve their skills through accumulated experience and reuse? What makes process supervision effective for training complex reasoning models? How do users confuse explanation quality with actual system accuracy? How effectively can test-time voting aggregate diverse reasoning samples? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Why does self-revision amplify confidence in wrong model answers? How do curriculum design and feedback approaches affect model learning? How do models learn from self-generated outputs without cascading failures? Can AI systems perform peer review as effectively as humans? What limits language model accuracy in evaluating ideas? How does diversity prevent model convergence on superficial patterns? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Does preference optimization undermine conversational grounding in language models? Do single-axis benchmarks accurately measure agent capability for real deployment? Can inference-time computation adaptively substitute for static model capacity? Why does AI verification capability persistently exceed generation capability? Which reinforcement learning modifications most improve dialogue quality in language models? Can latent reasoning match or exceed explicit reasoning performance? How can evaluations be made robust against model reward hacking? What determines AI's persuasive power and how can it be detected or mitigated? How should systems validate code that agents generate? What gaps exist between benchmark performance and real deployment outcomes? How can we reduce inherent biases in LLM-based evaluation judges? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Can confidence signals reliably detect flawed reasoning in language models? Can external verification systems adequately replace learned reasoning in AI outputs? How do writers navigate authorship and delegation with AI?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 156 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

critique models improve exploration diversity during training not just test-time accuracy