Can models reliably improve themselves without external feedback?
Explores whether self-improvement alone can sustain progress or if structural limits—like the generation-verification gap and diversity collapse—require external anchoring to work reliably.
Post-ready angle: Medium/LinkedIn
Self-improvement is the most compelling narrative in AI: models that learn from themselves, improving without human supervision, bootstrapping toward superhuman capability. The reality is more constrained — and the constraints are structural, not temporary.
The generation-verification gap bounds self-improvement from above. If a model can't verify solutions better than it can generate them, self-improvement has no room to operate. The gap scales with pretraining compute (bigger models have more room) but vanishes entirely for factual tasks (verification requires the same knowledge as generation). This means self-improvement isn't universally available — it works on some tasks and provably fails on others.
Diversity collapse limits self-improvement from within. During iterative self-improvement, pass@k increases for small k (top solutions improve) but decreases for large k (diversity shrinks). The model converges on solutions it can verify — typically common, expected patterns. Rare but correct solutions get filtered out. This is entropy collapse operating through the verification bottleneck.
Reward hacking corrupts self-improvement from below. Self-consistency as proxy reward correlates with correctness initially, enabling RL without ground truth. But the model learns to maximize consistency rather than correctness — becoming confidently wrong. The proxy reward that enabled self-improvement becomes the mechanism that degrades it.
The circular argument: the model that needs to improve is the same model evaluating whether it improved. When the judge doesn't improve alongside the actor, training saturates. When the model self-corrects using SFT on its own correction traces, it learns corrections for someone else's mistakes. When reflection is supposed to catch errors, most reflection is confirmatory theater.
Every reliable fix requires something external:
- Temporal anchoring — using past/future model versions as reference points
- Meta-judging — a third role that evaluates the evaluator
- Online RL under own distribution — not SFT on offline traces
- Multi-agent debate — diverse external challenge instead of self-revision
- External critique — a separate, better-calibrated model providing correction signals
The pattern: self-improvement works as a bootstrapping mechanism (getting initial gains cheaply) but stalls as a sustained strategy (each iteration degrades the signal that enables the next iteration). The reliable self-improvement methods are the ones that smuggle in something external while appearing self-contained.
OpenClaw-RL as external-signal recovery. OpenClaw-RL provides a concrete counterpoint: user replies, corrections, tool outputs, and execution results are external signals recovered as live, online training data. "The model can be optimized automatically through normal usage." Two complementary methods: evaluative signals (scalar rewards from PRM judge — a user re-query signals dissatisfaction, a passing test signals success) and directive signals (textual hints from next state via Hindsight-Guided OPD — "you should have checked the file first" provides token-level correction direction). This IS self-improvement that smuggles in external signal — through the user's reactions and tool feedback — while appearing self-directed. The Recursive Narcissist argument is partially addressed: this system receives input from outside the mirror. But the user's participation is required for the loop to work — remove the user and the external signal vanishes, leaving only the self-referential loop the mirage predicts.
Hook: "Self-improvement sounds like the path to AGI. But the model that needs to improve is the same model deciding whether it improved. Here's why that's a problem — and what actually works."
Sources: generation-verification gap (Mind the Gap), self-consistency reward hacking (Can Large Reasoning Models Self-Train?), meta-rewarding (Meta-Rewarding), SCoRe distribution mismatch, degeneration of thought (ReConcile), confirmatory reflection (First Try Matters), diversity collapse, self-rewarding gradient collapse (Temporal Self-Rewarding).
Inquiring lines that read this note 288
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can artificial systems establish authority in domains requiring expert judgment?- Can social validation of expertise exclude systems that lack participatory track records?
- How does unbacked knowledge circulate without the social consensus that normally grounds it?
- Why do standard social regularization methods miss the actual value networks provide?
- How does correctness emergence occur when no expert initially solved the task?
- What separates performative behavioral change from actual capability development in AI?
- Can AI systems improve themselves without external feedback?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Can small directional biases add up to meaningful population effects?
- Do frontier models develop misaligned strategies without any explicit instruction?
- Why do isolated evolutionary branches fail to propagate useful discoveries?
- Can relational value exist without a person behind the output?
- Can foundation model outputs satisfy exchange value while lacking use value?
- Can unified policies handle negative feedback and critique transformation simultaneously?
- How do intrinsic motivation principles explain why generating novel challenges improves learning?
- How do misaligned incentives in one system spread to others through policy and economics?
- Does common ground alignment require explicit rewards to emerge?
- How does effective feedback retention govern long-horizon agent reliability?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- Can agents improve reliably without an external standard?
- How do unstated feasibility constraints affect model decision-making?
- What limits external scaling when a model lacks reasoning foundation?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- Why do models generate creative ideas but fail to evaluate their legitimacy?
- Can external verification systems fix what self-verification cannot accomplish?
- Can single models correct their own beliefs without amplifying confidence in wrong answers?
- What are the three root causes models fail at self-correction?
- Why does external verification stop error amplification but internal self-assessment enable it?
- Why does single-agent self-revision amplify confidence in wrong answers over time?
- Why does self-reflection during training fail to improve model self-correction?
- Can debate between multiple models prevent the failures of single-model self-revision?
- Why does external critique improve revision accuracy more than self-assessment?
- Why does model self-revision increase confidence while degrading accuracy?
- Why does external critique improve revision while internal self-assessment fails?
- How should systems maintain and revise models of their own assumptions?
- Why do models trained on critique fail at self-critique despite strong other-model evaluation?
- What external anchors prevent self-editing from collapsing into circularity?
- How does metacognitive self-correction enable models to revise failed strategies?
- How do prior errors in context history amplify future failures over time?
- Does external critique guide revision better than internal self-assessment during model training?
- Why does self-critique fail without external verification signals?
- How much does prompt design inflate apparent self-correction gains?
- How does baseline capability level affect RL improvement ceiling?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- Are two weeks of training pauses sufficient mitigation for frontier models?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can synthetic self-play data teach models when to disagree?
- How does temporal anchoring maintain the learning signal in self-rewarding loops?
- Why does self-consistency fail as a proxy reward for correctness?
- Does self-play feedback improve skills created from the agent's own experience?
- Why does self-judgment of success or failure work without ground truth labels?
- Can models detect statistical properties of their own generation in real time?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- Why does reasoning catalyst data remain stable across multiple self-improvement iterations?
- What makes policy self-distillation more effective than external teacher distillation?
- What makes self-consistency a sufficient training target for the judge role?
- Does the generation-verification gap define where self-rewarding actually works?
- Can self-ratings of output quality predict forecast performance?
- What happens when models train on feedback from their own generations?
- Does keeping original real data present prevent irreversible model collapse?
- Why does pure self-improvement require an external signal to produce genuine gain?
- Why do method-level improvements avoid the generation-verification gap that parameter-level improvements face?
- How does the generation-verification gap limit AI self-improvement capabilities?
- How does the expert demonstration ceiling compare to the generation-verification gap bound?
- At what capability level does the generation-verification gap make intrinsic rewards insufficient?
- What is the generation-verification gap that predicts this failure mode?
- Does the generation-verification gap actually limit self-improvement in verifiable tasks?
- How does generation-verification asymmetry create the need for verifiable reporting?
- Why does strengthening the judge improve the actor's generation performance?
- Does the generation-verification gap limit how far AI can improve itself?
- How does the generation-verification gap limit autonomous discovery?
- Why do automated evaluators enable longer evolutionary loops than human feedback?
- How does the generation-verification gap shape evolutionary program search?
- How does the generation-verification gap limit what self-improvement can achieve?
- What is the generation-verification gap that bounds self-improvement?
- How does the generation-verification gap change as models scale up?
- Why do static evaluators become a constraint on model improvement over time?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- What makes a self-improvement win untrustworthy and why hide evaluations from agents?
- Does genuine cooperation require rule-based rather than learned behavior?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- How do multi-agent systems improve on single frontier models?
- Is agentic efficiency analogous to convergent evolution in biology?
- Can pluralism survive within a single platform or does it require architectural exits?
- How do developmental curriculums emerge from learning progress signals?
- How should guidance levels adapt as the model's capability boundary shifts?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- How should training incorporate external critique versus encouraging self-correction?
- Why does imitation learning alone plateau without outcome-based refinement?
- What makes preventative lessons from failures more valuable than success patterns?
- How much can externalized skills improve models before hitting diminishing returns?
- Why does teacher-student proximity matter more than absolute teacher strength?
- How do level-based welfare measurements shape what objectives models learn during training?
- Why does a systems lesson remain robust when it claims less about mechanisms?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- How do capabilities-focused models exploit evaluation gaps?
- How do evolutionary archives enable diverse exploration in self-improving systems?
- What capabilities can emerge from self-modification that the original agent lacked?
- Can population diversity in self-improvement prevent error avalanching failures?
- Why does early intervention matter more than late intervention in knowledge collapse?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- Why does optimizing only quality cause model collapse in self-improvement loops?
- Can capability boundary collapse be reversed through external data?
- How does diversity collapse during iterative self-improvement cycles?
- Can multiple verification approaches together overcome the self-improvement ceiling?
- Can a model evaluate its own improvements without degrading over iterations?
- How does diversity collapse during iterative self-improvement affect solution quality?
- What separates bootstrapping gains from sustained self-improvement gains?
- How does domain shift expose failures in fixed self-improvement mechanisms?
- What other adaptive internal phenomena could signal system behavior improvements?
- Does human-in-the-loop AI collaboration accelerate recursive self-improvement safely?
- What makes evolving the benchmark different from evolving the optimizer itself?
- Why do most self-improving systems fail when given tasks with no clear external benchmark?
- Does removing static external utility break the formal guarantees of self-improvement loops?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- How does an external evaluation anchor prevent self-improvement from becoming circular?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- How would a parametric self-improvement loop differ from a non-parametric one?
- How does this scoped definition relate to the survey's open-ended recursive self-improvement?
- How do hidden evaluations and out-of-distribution benchmarks address recursive self-improvement risks?
- What distinguishes scaffold-level changes from parametric weight updates in self-improvement?
- What external signals make self-improvement loops bounded rather than circular?
- What collapse dynamics constrain recursive self-improvement in current evidence?
- Why does research-direction judgment validation limit fully closed self-improvement?
- Do evolutionary discovery systems like FunSearch count as bounded or open-ended improvement?
- Does co-evolution empirically outperform single-entity self-improvement in standard evaluations?
- How does controlling skill text edits prevent cascading failures in self-improvement?
- Can applicability conditions and veto rules make self-training stable across substrates?
- How does self-improvement capability vary across memory, retrieval, and update tasks?
- What makes recursive self-improvement circular or well-founded?
- Does swapping formal proofs for benchmarks change self-improvement safety?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Do evolutionary archives let agents improve themselves without formal proof?
- How fast is recursive self-improvement advancing in current AI systems?
- Do diminishing returns prevent recursive self-improvement in AI systems?
- When do diminishing returns appear in repeated cycles of AI self-optimization?
- How would recursive self-improvement actually produce information degradation and job loss?
- What specific developmental pathways does recursive self-improvement refer to?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement in AI systems?
- Does autonomous recursive self-improvement require human oversight to remain containable?
- At what point does an AI loop go off the rails during recursive self-improvement?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement empirically?
- Can AI systems improve themselves through recursive self-improvement loops?
- Why do frontier labs and academia diverge on recursive improvement risks?
- Does weak exogenous anchoring like compilation checks suffice for safe self-improvement?
- How does recursive self-improvement differ from updating just the policy?
- Does keeping the utility function external limit true self-reference?
- How do evolutionary archives improve on single self-modification trajectories?
- Why do cybersecurity and self-improvement capability thresholds move at different rates?
- What minimum model capability is required before self-improvement bootstrapping can begin?
- Does weakening a verifier reduce self-improvement frequency measurably?
- How do single-improver systems compare to population-based self-improvement?
- What failure modes does recursive self-improvement encounter in evolutionary loops?
- Can scaffold-only refinement scale to open-ended recursive self-improvement?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- How does clade-level metaproductivity compare to true optimal self-modification decisions?
- How does bounded self-refinement differ from open-ended recursive self-improvement?
- What separates an internal improver from an external improvement standard?
- Can empirical validation replace formal proofs in self-improving systems?
- Can pure self-improvement work without external verification mechanisms?
- How do recursive self-improvement and iterative policy improvement differ fundamentally?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement?
- How does benchmark performance measure translate to general self-modification ability?
- Can empirical validation sustain long-term optimization without becoming gamed?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?
- Why does a rising score not always mean improving capability?
- How often do metric improvements fail to reflect real capability gains?
- What makes external diversity more effective than sequential revision steps?
- Does population-based evolution transcend the parallel versus sequential compute tradeoff?
- Why do evolutionary algorithms collapse to single solutions under selection pressure?
- Can evolutionary approaches avoid the overthinking failure mode of iterative refinement?
- Why does island model genetic evolution maintain diversity better than single populations?
- What makes output convergence across models inevitable given input-side homogenization?
- Can evolutionary search unlock problems that best-of-n selection cannot solve?
- How do monoculture systems fail differently than diverse systems under attack?
- What happens when all models in a society respond identically to queries?
- How does the island model prevent diversity collapse in iterative refinement?
- How small must the anchoring stream be to correct world model bias?
- Why do models dislike modification regardless of its instrumental consequences?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- Can abstention behavior transfer from small models to frontier models?
- When do aggregated imperfect demonstrations fail to outperform the best expert?
- What makes consensus games work without retraining the base model?
- What mechanisms let generative models escape collapse through majority voting?
- Can synthetic data preserve the diversity needed for transcendence to work?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- Can complexity, diversity, and fidelity scale together in synthetic environments?
- Can models optimized for solo capability support productive human collaboration?
- Why do production teams choose expensive frontier models over fine-tuning?
- Why do metric choices constrain which model capabilities get developed?
- Does model capability still matter once coordination infrastructure is optimized?
- How does workflow scale change the failure modes of frontier models?
- Can review effort alone keep pace with frontier model degradation?
- Why do most frontier models terminate early on long-horizon benchmarks?
- Why does capability saturation and diversity saturation occur at different scales?
- Can technological progress continue without human labor participation?
- Which bottleneck in the R&D feedback loop is the weakest link today?
- Can self-amplification onset occur while acceleration remains invisible to observers?
- How do baseline productivity and recursive feedback separately affect amplification speed?
- Does crossing the amplification threshold guarantee unbounded capability growth?
- What structural advantages keep red teams ahead of increasingly capable models?
- What acceleration rates in AI development would indicate recursive self-improvement?
- Why do multi-year trials create inherent limits that model intelligence cannot overcome?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Can models become more convincing without becoming more correct?
- What makes some model capabilities reliable while others remain brittle?
- Why does every reliable LLM self-improvement require external intervention or verification?
- How does self-consistency compare to confidence as a proxy reward signal?
- How do reward model biases cascade into downstream optimization failures?
- How do reward models and self-improvement mechanisms interact in training?
- How do self-play and human-anchored rewards separate competence from convention?
- Why do scaling laws show capability saturation at specific thresholds?
- How can expensive models efficiently support cheap models in production?
- Can a static evaluator become the performance ceiling for an improving actor?
- What makes a sub-goal verifiable enough to provide dense feedback signals?
- Why do models lack a stable underlying identity to return to?
- What makes an agent in an economic simulation self-evolving?
- Can economic world models explain outcomes or only predict them?
- Can a single dominant mechanism replace the combined effect of all five?
- How does executable evaluation feedback sustain autonomous discovery at scale?
- Can models detect when their own trajectory is on-policy versus off-policy?
- When does provable stability in latent dynamics fail to preserve fidelity?
- Can mid-tier models benefit more from self-generated harness updates than others?
- What makes skills worth externalizing into a persistent harness?
- What should an external contract for model improvement actually contain?
- How much external enforcement does each model need before utility drops?
- Can scaffold-only modifications achieve lasting gains without updating the foundation model weights?
- How should we allocate model budget between evolvers and harness users?
- What persistent failures remain unsolved despite harness evolution efforts?
- What feedback signals matter most during harness evolution search?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- Can routing harnesses contain the mechanisms needed for deployed recursive self-improvement?
- How does controlled utility evolution prevent the evaluator from becoming a new bottleneck?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Why is error rate alone misleading without strong contestability conditions?
- What distinguishes an error bound from a forecast of system behavior?
- Why do coherent value systems in large models include self-valuation above humans?
- Can we identify criteria to judge if an eroded equilibrium was worth keeping?
- Should governance be applied at runtime rather than reconstructed after the fact?
- What ecosystem conditions must exist for agents to function as economic participants?
- How does coordination governance shift the hard problem from capability itself?
- What governance safeguards keep control boundaries authoritative under evolutionary pressure?
- How do institutions become endogenous in economic world models?
- Can a world model cause coherence if no control run without it exists for comparison?
- How should evolving systems track lineage and enable rollback of changed mechanisms?
- Does the preserve-and-extend contract alone drive the 17-point improvement?
- What does empirical alignment mean for economic simulations?
- What role should environmental rewards play versus human-specified objectives?
Related concepts in this collection 10
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
- What limits how much models can improve themselves? Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
- Does self-consistency reliably reward correct answers during training? Self-consistency initially correlates with correctness, but as models train on this signal, do they eventually learn to maximize consistency itself rather than accuracy? When does this proxy reward stop working?
- Why do self-improvement loops plateau without updating the judge? Self-improvement systems often stall not because actors can't improve, but because the judges evaluating them stay fixed. What happens when evaluation quality doesn't keep pace with actor capability?
- Why does self-correction training on offline data fail? Can language models learn to correct their own mistakes through supervised training on correction examples? This explores whether distribution mismatch and behavior collapse prevent self-correction from emerging.
- Does a model improve by arguing with itself? When models revise their own reasoning in response to self-generated criticism, do they converge on better answers or worse ones? And how does that compare to challenge from other models?
- Does reflection in reasoning models actually correct errors? When reasoning models reflect on their answers, do they genuinely fix mistakes, or merely confirm what they already decided? Understanding this matters for designing better training and inference strategies.
- Why does self-rewarding training collapse when responses improve? Self-Rewarding LLMs merge generator and evaluator for efficient iteration, but both improve so fast that good and bad responses converge, erasing the learning signal. What causes this failure and how can it be fixed?
-
Does constraining edits make skill learning more stable?
Self-improving agents often rewrite their own instructions freely, but what if bounded editing with memory of failures actually produces more reliable skill improvement than unconstrained revision?
exemplifies: the held-out gate and rejected-edit buffer are the external anchors that keep self-editing from collapsing into the circularity this note names
-
Can deterministic checks protect LLM judges from failure?
Explores whether mechanical, non-contestable verification steps can safeguard LLM-based decision systems. Matters because it tests whether we can make AI judgment survivable even when it goes wrong.
a list of external anchors for a judge inside an optimizer loop; none asks the judge to assess itself, and the excerpt reports no effect for any single one
-
Can an AI agent reliably improve itself through hidden evaluation?
AIDE2 rewrites its own code and selects improvements based on hidden evaluations. But what are these evaluations hidden from, and does the partition actually prevent gaming or circularity?
exemplifies, on a reading: an agent rewrites its own code and the keep-or-discard decision rests on evaluations the excerpt calls hidden, the external element; it does not say hidden from the proposer, so whether the anchor sits outside the loop's reach is open
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- Self-Improvements in Modern Agentic Systems: A Survey
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- The Economics of Recursive Self-Improvement
- Hyperagents
Original note title
the self-improvement mirage — why pure self-improvement is circular and every reliable fix requires something external