INQUIRING LINE

When an AI gain lasts longer than similar studies found, one thing to ask is whether it measured real skill or memorized test answers.

Why did the PR lift persist longer than in comparable studies?

This reads 'PR lift' as a measured improvement (possibly in pull-request output, or in some 'PR'-named metric) that lasted longer than similar studies found, and asks what explains that durability. None of the retrieved notes reports this result, so the answer below covers what the corpus says about why a measured gain lasts or fades.


This reads 'PR lift' as a measured improvement (possibly in pull-request output, or in a metric abbreviated PR) that lasted longer than similar studies found, and asks why. None of the twelve retrieved notes describes this study, so I can't explain that particular result. 'PR' is also unexpanded, which makes the question hard to match. What the notes do offer is a way to judge any claim that a gain persisted: lasting gains and quickly fading gains usually come from different mechanisms.

The first thing to check is what the lift was actually measuring. In the RLVR work, models seem to improve on math benchmarks, but much of that gain is memorized test data. One model can reconstruct over half of a well-known test set from partial prompts, yet scores zero on problems released after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. A gain like that can look perfectly stable while measuring nothing real. A related note shows that real skill gains and inflated benchmark scores can happen at the same time Can genuine reasoning activation coexist with contaminated benchmarks?. A durable score doesn't by itself mean a durable skill.

The second check is whether the gain came through the route the study claimed. In PAST-Bench, agents with the same overall improvement turned out to reach it in different ways. Some actually used their saved experience; others got there without it Do agents with the same performance gain follow the same learning pathway?. If a study's lift lasts longer than its peers', the mechanism may differ from the peers' too. Only direct evidence of the mechanism can tell you that, not the score.

The corpus also shows how gains usually fade. AI persuasion starts strong and then weakens over repeated rounds with the same person, while human persuaders stay steady Does AI persuasiveness fade across repeated conversations with the same person?. Self-improvement stalls unless something outside the model keeps it honest, such as a judge, user corrections, or tool feedback Can models reliably improve themselves without external feedback?. A cohort of different peer models gives that kind of outside signal more reliably than a model grading itself Can peer models replace external judges for reward signals?. Gains also tend to last when they sit in the system around the model rather than in the model. Better execution harnesses carried over to newer models with no changes Can execution harnesses lift model performance without retuning weights?.

So a lift that persists usually points to one of three things: the measurement is contaminated, the source of improvement is external and keeps renewing, or the gain lives in a reusable system. Durability also needs the right data to show it. One self-improvement study reported seven successive gains but no score trajectory, so it couldn't show whether returns held up or shrank Does recursive self-improvement sustain gains or hit diminishing returns?. If you can say which study 'PR lift' comes from, these three checks give you a framework for questioning it.


Sources 8 notes

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Can genuine reasoning activation coexist with contaminated benchmarks?

RLVR activates genuine reasoning patterns through RL training while benchmark improvements may reflect data memorization on contaminated datasets. These operate at different measurement levels and can coexist without contradiction.

Do agents with the same performance gain follow the same learning pathway?

PAST-Bench testing reveals that agents can achieve the same overall improvement through different mechanisms—some genuinely using saved experience while others do not. Mechanism evidence paired with performance scores is needed to distinguish true learning from apparent gain.

Does AI persuasiveness fade across repeated conversations with the same person?

Claude and DeepSeek showed strong initial persuasive advantage, but this edge eroded across repeated quiz rounds while human persuaders maintained consistent effectiveness. This decay pattern is opposite to human-to-human persuasion, where rapport typically strengthens over time.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Show all 8 sources
Can peer models replace external judges for reward signals?

Co-RL trains decoupled models using peer predictions as rewards, avoiding the bias and collapse of self-generated feedback. Heterogeneous cohorts consistently improve reasoning across benchmarks and often match ground-truth supervised training.

Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Does recursive self-improvement sustain gains or hit diminishing returns?

The paper reports seven successive improvements in an 8-day run but provides neither the magnitude of each gain nor their timing. Without score trajectories and longer-horizon data, the evidence supports only that improvements transferred, not that recursive self-improvement sustains returns against diminishing curves.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.