INQUIRING LINE

Does a Kaggle medal earned in the past year tell you more about someone's next result than an older one?

How much does medal age matter when predicting Kaggle performance?

This explores whether the age of a Kaggle medal (how long ago someone earned it) changes how well it predicts their performance in a new competition, and whether that changed once generative AI arrived.


This explores whether a Kaggle medal earned recently tells you more about someone's next result than one earned years ago, and whether AI tools changed that. The short answer: medal age matters a great deal. A study of 444,698 competition entries found that medals predicted performance on Kaggle's hidden test sets almost entirely during their first year Do Kaggle medals still predict performance after AI arrived?. After that, an old medal adds little. This pattern held both before and after generative AI arrived. You might expect AI assistance to make every credential worthless, but fresh medals kept most of their predictive value. So the useful question is how recently a medal was earned, not just whether someone has one.

The less obvious finding is about what happens when a medal outlives the system that awarded it. Kaggle retired one competition format, where participants uploaded prediction files, before the AI era began. Those medals then aged on their normal schedule, but profiles kept displaying them at full value. The same audit attributes about half of the drop in how well upload-format medals predicted performance to this "institutional stranding" How much did retiring a competition format hurt medal credibility?. The platform stopped running the checks behind the badge, but the badge stayed up. That means some of what looks like "AI broke the credentials" is really "the platform let a credential drift away from the process that verified it."

The same idea shows up elsewhere in the collection: a score only means something under the conditions it was earned in. Language models show a similar effect with benchmarks. A model can rebuild more than half of an older math benchmark from memory yet score zero on problems released after its training, so the old score reflected exposure to the test rather than skill Does RLVR success on math benchmarks reflect genuine reasoning improvement?. An old medal and a contaminated benchmark score fail the same way: the number stays put while its link to current ability weakens.

The collection is thin here. Only two notes deal directly with Kaggle medals, and they come from the same audit. It doesn't say how quickly value fades within that first year, whether it differs by competition type beyond the retired format, or how individual competitors' skills change over time. To go deeper, start with the first note for the overall decay pattern, then read the stranding note for the less expected point: how the platform handles old badges matters about as much as the arrival of AI.


Sources 3 notes

Do Kaggle medals still predict performance after AI arrived?

Across 444,698 participations, medals predicted hidden-test performance almost entirely through their first year in both pre- and post-AI eras. Fresh medals retained most value after generative AI arrived, suggesting verified credentials stayed informative despite platform changes.

How much did retiring a competition format hurt medal credibility?

The audit attributes roughly half the decline in upload-format medal informativeness to institutional stranding: the platform retired the format before AI, medals aged on schedule, yet stayed visible at their original value. This decoupled the credential from the validation mechanism it once represented.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.