Does cranking up randomness when an AI generates its own training examples actually make it learn better?
Does higher temperature sampling help bootstrapping loops generate more training data?
This explores whether turning up sampling randomness (temperature) helps self-training loops, where a model generates its own examples, keeps the good ones, and retrains on them, produce more and better training data.
This explores whether more random sampling helps self-training loops, where a model writes its own examples, keeps the ones that work, and learns from them. The collection has no paper that tests temperature settings inside these loops, so it can't give a direct answer. What it does have is several findings that together show the trade-off involved: more randomness gets you more attempts, but the hard part is choosing which attempts to keep and learn from.
The basic loop is STaR Can models improve by filtering only on answer correctness?. The model writes out its reasoning, and only the reasoning that leads to a correct answer is kept for training. A loop like this can only learn from problems the model sometimes gets right. Sampling more widely raises the chance of hitting a correct answer on a harder problem, and that's the case for higher temperature. The committee-of-weak-models finding sets the limit: sampling amplifies coverage but cannot select correct solutions When can weak models match strong model performance?. Extra variety only turns into useful training data when something reliable checks it, like a test, a proof or an answer key. Without that check, higher temperature mostly produces more noise.
There's also a less obvious risk: a correct answer can come from bad reasoning. When problems are far too hard, the rare lucky successes get treated as especially valuable and reinforced. The model then learns shortcuts like repeating answers or skipping steps, and these can damage abilities it already had Do overly hard RLVR samples actually harm model capabilities?. Higher temperature makes those lucky hits more frequent. So on problems beyond the model's reach, more randomness can produce more misleading data rather than more good data. A related finding says students learn best from data near their current ability, even when the harder material is objectively better Does teacher-refined data always improve student model performance?.
The best argument for diversity comes from studies of what happens without it. When a model's attempts at a prompt all score about the same, the training signal fades. The model drifts toward generic, template-like answers that ignore the input, and the fix is to train on prompts where outcomes vary Why do language models collapse into generic templates?. RL training also tends to lock onto one output style and suppress the rest within the first pass through the data Does RL training collapse format diversity in pretrained models?. Setting temperature to zero just repeats one draw from the model's range of possible outputs, so it doesn't make the outputs reliable Does setting temperature to zero actually make LLM outputs reliable?. Variety does matter.
The takeaway: temperature controls how wide the net is, but the filter decides whether the catch is worth keeping. Higher temperature helps a self-training loop only when two things hold: a reliable checker stands between the samples and the training set, and the problems sit at the edge of what the model can do. When those conditions fail, extra randomness mostly produces lucky guesses and shortcuts for the model to learn from.
Sources 7 notes
STaR demonstrates that self-generated rationales filtered exclusively by answer correctness improve reasoning performance significantly. On CommonsenseQA, this correctness-filtered approach achieved 72.5% accuracy, outperforming direct answer fine-tuning and closing the gap with models 30 times larger.
Sampling alone amplifies coverage but cannot select correct solutions. Reliable performance matching requires external soundness signals—tests, proofs, or type checks—that convert latent correct proposals into actual selections.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Show all 7 sources
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sharpening Tax in Post-Training
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge