Can models improve by filtering only on answer correctness?
Does a bootstrapping loop that selects self-generated rationales based solely on whether the final answer is correct—without evaluating reasoning quality—actually improve model performance on reasoning tasks?
STaR (the "Self-Taught Reasoner") proposes a loop: few-shot prompt a pretrained LLM M to generate a rationale and answer for each problem in a dataset D of (question, answer) pairs; keep only the rationales whose final answer matches the ground truth ("we assume that rationales that lead to correct answers are of better quality than those that lead to incorrect answers"); fine-tune the original pretrained model on that filtered set; and repeat, regenerating from the newly fine-tuned model each round. The check that decides what counts as good training data is exclusively whether the answer came out right, never an evaluation of the rationale's content. On CommonsenseQA this produced a +35.9% gain over a few-shot baseline and +12.5% over a model fine-tuned to predict answers directly, reaching 72.5% versus 73.0% for a fine-tuned model 30 times larger.
The loop has a built-in failure mode the paper names directly, "this loop eventually fails to solve any new problems in the training set because it receives no direct training signal for problems it fails to solve." Their fix, "rationalization," hands the model the correct answer as a hint and asks it to generate a rationale backward from that answer — sampling from p(r | x, y) instead of p(r | x) — then strips the hint before adding the rationale to the fine-tuning set, "as if the model had come up with the rationale without any hint." The paper frames this as a pipeline-level loop, not a model editing its own weights mid-inference: "we propose what is, to our knowledge, the first technique to allow a pre-trained large language model to iteratively use its language modeling capacity to improve itself," and each round retrains from the original M rather than continuing to train one model, specifically to avoid compounding overfitting. The authors also tested the obvious alternative to rationalization, higher-temperature sampling to get more training examples, and found it "counterproductive": it raises the odds of a correct answer reached by bad reasoning, and training on that reasoning "prevents generalization," with scratchpads in arithmetic tasks "diverg[ing] into meaninglessness."
This is the correctness-filtered bootstrapping loop that Can models learn reasoning from predicting any text? extends from curated question-answer datasets to arbitrary token positions in ordinary text — STaR itself only works where D already supplies a ground-truth answer to filter against. It also names the condition Can models improve themselves on tasks without verifiable answers? is built to escape: STaR's own limitations section concedes the method depends on domains with checkable final answers, which is exactly the constraint later catalyst-data approaches try to relax for open-ended instruction-following. And STaR's requirement that the base model already clear "above chance" few-shot performance before bootstrapping can start — "we found that GPT-2 was not able to bootstrap from few-shot reasoning in even the arithmetic domain" — is the same precondition described more generally in Do base models already contain hidden reasoning ability?.
The paper does not show the loop compounding indefinitely — it reports running the method "until the performance plateaus," not a result where successive iterations keep climbing, and the authors flag that in domains with high chance performance (e.g., binary decisions) "many poor rationales" confound the filter, calling the fix "an open problem." Because the only check in the loop is final-answer correctness, the method is structurally blind to a rationale that reaches the right answer for the wrong reason — a risk the paper surfaces in its temperature experiments but does not fully resolve for the rationalization pathway, where nothing outside the dataset's own answer key confirms that a backward-generated justification is a sound argument rather than a merely plausible-sounding one.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can reasoning traces reveal actual model reasoning versus plausible output? How do models learn from self-generated outputs without cascading failures? How do curriculum design and feedback approaches affect model learning? What limits recursive self-improvement in autonomous AI systems? How can evaluations be made robust against model reward hacking? Why does polished AI output gain credibility despite fundamental verifiability problems?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can models learn reasoning from predicting any text?
Does training rationale generation at every token position on arbitrary internet text enable general reasoning without task-specific supervision? This challenges the assumption that reasoning requires curated QA datasets.
extends STaR's correctness-filtered loop from curated QA datasets to token-level rationale generation over arbitrary text
-
Can models improve themselves on tasks without verifiable answers?
Most self-improvement methods require verifiable correctness signals like math or code. Can models improve on open-ended instruction tasks where right answers aren't automatically checkable? And what minimal training is needed to unlock this?
relaxes the checkable-answer constraint STaR's own limitations section concedes
-
Do base models already contain hidden reasoning ability?
Explores whether reasoning capability emerges during pre-training as a latent feature rather than being created by post-training methods like reinforcement learning or fine-tuning.
matches STaR's finding that bootstrapping requires a base model already above chance on few-shot reasoning
-
When does explicit reasoning actually help model performance?
Explicit reasoning improves some tasks but hurts others. What determines whether step-by-step reasoning chains are beneficial or harmful for a given problem?
Qualifies STaR's bootstrapping: explicit reasoning it reinforces helps logical-derivation tasks but degrades continuous-judgment tasks
-
Can models that reason well also grade reasoning well?
Do the same capabilities that let language models produce valid reasoning also let them spot flawed reasoning? Testing this assumption reveals a surprising gap between production and evaluation skills.
Evidence for STaR's risk: outcome-only filtering builds production but leaves verification of flawed-but-correct reasoning underdeveloped
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- STaR: Bootstrapping Reasoning With Reasoning
- Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
- Large Language Models Cannot Self-Correct Reasoning Yet
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
- Do Reasoning Representations Help Humans Evaluate LLM Outputs?
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Reverse Thinking Makes LLMs Stronger Reasoners
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
Original note title
STaR bootstraps reasoning by filtering self-generated rationales on answer correctness — not rationale quality