Do critique models improve diversity during training itself?
Explores whether critique integrated into the training loop, beyond test-time scoring, actively maintains solution diversity and prevents the model from converging too narrowly during iterative self-training.
The intuitive framing of critique models is that they help at test time: the model generates, the critic scores, we select the best. But the more important finding from AutoMathCritique is that critique integrated into the training loop improves the actor model's exploration efficiency and solution diversity during training itself.
Without critique in the loop, iterative self-training suffers from "tail narrowing" — the model converges on a narrow distribution of solutions, becoming less able to explore diverse reasoning paths. The critique model counteracts this: by providing step-level feedback on exploration, it guides the actor toward high-quality paths it wouldn't have discovered alone, maintaining distributional breadth through training.
This connects to Does policy entropy collapse limit reasoning performance in RL?: critique models are a way to maintain entropy — the exploration needed for continued improvement — without relying solely on architectural entropy management (Clip-Cov, KL-Cov). The critique is an external signal that prevents premature convergence.
The implication: critique models are training infrastructure as much as inference infrastructure. Evaluating them only on test-time accuracy misses their more fundamental role.
Inquiring lines that read this note 103
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do agents learn to distinguish valuable feedback from noise?- Can unified policies handle negative feedback and critique transformation simultaneously?
- How do intrinsic motivation principles explain why generating novel challenges improves learning?
- Why does embedding evaluation criteria in prompts reduce creative scope?
- Can runtime interventions like meta-cognitive prompting work where training interventions fail?
- Does diversity prompting actually help models explore human argument space?
- Why does grouping feedback by shared correction pattern improve prompt generalization?
- Can few-shot examples narrow generative diversity in creative tasks?
- Can prompting for specific creative paradigms improve ideation diversity?
- Why do research ideation systems suffer from diversity collapse despite high novelty metrics?
- Can diverse human creativity survive if all AI systems converge on similar outputs?
- What happens to idea diversity when AI tools draw from collective knowledge?
- How can semantic diversity optimization work if exploration and exploitation were truly opposed?
- How does directional diversity compare to other forms of parallel planning?
- Can LLM diversity collapse in research ideation be reversed or mitigated?
- What makes creative writing diversity different from code diversity fundamentally?
- Which aggregation method best exploits diversity in generated solutions?
- Does AI refinement preserve idea diversity better than AI ideation compresses it?
- What would a valid diversity measure for AI-assisted ideation tasks look like?
- How does critique fine-tuning on one problem unlock broader reasoning?
- Can diverse critiques on a single problem unlock reasoning without diverse problem sets?
- Can population diversity in self-improvement prevent error avalanching failures?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- How does diversity collapse during iterative self-improvement cycles?
- How does diversity collapse during iterative self-improvement affect solution quality?
- Does co-evolution empirically outperform single-entity self-improvement in standard evaluations?
- What makes external diversity more effective than sequential revision steps?
- How does an aggregator use diverse complementary traces to improve final answers?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- How does forced exploration through diversity rewards differ from suppression-based negative reinforcement?
- How do semantic reward shaping approaches compare to full critique models?
- How do you verify whether your context distribution satisfies covariate diversity?
- What conditions make training diversity better than individual expert quality?
- Why does positive reinforcement degrade diversity at higher k values?
- What creates the irreducible trade-off between quality and diversity in training data?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- How much does diversity training cost in single-shot pass@1 performance?
- Does verbalized sampling preserve factual accuracy and safety during diversity gains?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- How do complexity and diversity affect model performance differently?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Why does diversity of training cases matter more than raw dataset size?
- How do ensemble methods apply within a single model?
- Can structural diversity through role assignment replace emergent diversity in small models?
- Why does evaluating multiple candidates work better than judging one answer?
- How does majority voting fail when reasoning samples lack genuine diversity?
- Can explicit rejection responses solve the over-specialization failure mode?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- Can rejected edits serve as negative feedback like hard negatives in contrastive learning?
- Can debate between multiple models prevent the failures of single-model self-revision?
- Why does external critique improve revision accuracy more than self-assessment?
- How does symbolic solver feedback differ from language-based self-critique?
- Why does external critique improve revision while internal self-assessment fails?
- Why do models trained on critique fail at self-critique despite strong other-model evaluation?
- Can external retrieval signals outperform internal self-assessment during revision?
- Does external critique guide revision better than internal self-assessment during model training?
- Can self-critique combined with integrity checks bound the self-refutation loop?
- How should training incorporate external critique versus encouraging self-correction?
- Why does critique training produce deeper understanding than imitation training?
- Does critique training improve exploration diversity during model training or only test time?
- Does the productive difficulty band ever stabilize during training?
- Does training on curated solutions transfer to unseen problem types?
- Does self-generated training data reduce a model's capability diversity?
- Can self-training drift be prevented by applying student compatibility filtering?
- What makes self-consistency a sufficient training target for the judge role?
- What happens when models train on feedback from their own generations?
- What causes code quality to degrade across multiple rounds of recursive self-training?
- Can shifting the accuracy metric itself eliminate the need for diversity post-processing?
- How does the island model prevent diversity collapse in iterative refinement?
- Do draft-and-revise loops work better when guided by unresolved constraints than by diffusion-style denoising?
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Why does outcome-based RL specifically lose diversity during training?
- Why does preference tuning reduce diversity in code but increase it in creative tasks?
- What happens to model grounding when preference optimization increases effective diversity?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- Should test-time search maximize diversity of competent solutions instead of converging on one strategy?
- Why does test-time search also prioritize diversity over single-best convergence?
- Does semantic diversity in output space compete with reward-component diversity?
- Can diversity-aware reward bonuses achieve what set-level objectives achieve naturally?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
critique models as a mechanism against entropy collapse
-
Can natural language feedback overcome numerical reward plateaus?
Exploring whether chain-of-thought critiques can push past performance ceilings that scaling data alone cannot break in reinforcement learning for reasoning tasks.
concrete evidence: Critique-GRPO shows that CoT critiques break plateaus where 8x scaling of numerical rewards fails; the NLF mechanism works precisely because critiques expand the effective exploration space that numerical rewards cannot reach
-
Can diversity optimization improve quality during language model training?
Standard RL training assumes quality and diversity trade off, with diversity optimization potentially hurting performance. Does explicitly rewarding semantic diversity during reinforcement learning actually improve output quality alongside diversity?
DARLING provides the complementary mechanism: critique models maintain diversity by guiding exploration quality, while explicit semantic diversity optimization maintains diversity by directly rewarding distributional breadth — together they address the entropy collapse problem from both the feedback channel (critique) and the reward signal (diversity bonus)
-
Can a single problem unlock reasoning through solution critique?
Does exposing models to diverse critiques of different solutions to one problem activate reasoning as effectively as training on many problems? This tests whether solution diversity matters more than problem diversity.
extends with extreme efficiency: CFT shows that diverse critiques on a *single* problem suffice for reasoning activation — the diversity-via-critique mechanism does not need a diverse problem distribution, only diverse critiques of the solution space; this is the strongest evidence for the "critique is training infrastructure" framing
-
Does critiquing errors teach deeper understanding than imitating correct answers?
Can training models to critique flawed responses build better structural understanding than standard supervised fine-tuning on correct answers? This matters because it reveals whether deep reasoning requires engaging with failure modes rather than pattern matching.
extends to the training data design: training models on critiques of noisy responses produces deeper understanding than training on correct responses; the principle generalizes from "critique guides exploration" to "critique IS the training signal"
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
- Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
- Outcome-based Exploration for LLM Reasoning
- Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
- Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration
- Jointly Reinforcing Diversity and Quality in Language Model Generations
Original note title
critique models improve exploration diversity during training not just test-time accuracy