Stranded Credentials: Keeping Online Reputation Systems Informative in the AI Era
Abstract Platforms summarize providers’ past achievements into credentials that buyers use to judge quality. Generative AI can now produce much of the work those achievements certify, raising fears that those quality signals are worthless. We audit how well such credentials stay informative in Kaggle’s 2010–2026 archive, where medals are won on predictions scored against withheld answers and two evaluation formats ran side by side. Across 444,698 participations, a medal’s power to predict performance sits almost entirely in its first year, in both formats and eras. Fresh medals kept most of their value through the AI transition. About half of the collapse in the informativeness of one format’s medals is institutional: the platform had been phasing out that format before AI, and its medal stock aged out on schedule. Old medals look more informative only in isolation. The platform’s official lifetime-tier display discards up to a sixth of the medals’ predictive power. A recency-weighted index fit before the AI era explains AI-era performance about 13% better than the tiers and selects entrants who perform better on average, though the tiers still identify extreme top performers better. Displaying a recent-performance summary alongside the lifetime tiers would recover the discarded information for buyers.
Introduction. Platforms that summarize providers’ past achievements into credentials must keep those credentials informative for buyers as AI changes how work is produced. Freelance marketplaces award badges for highly rated deliverables, review platforms rank reviewers by the quality of their past reviews, and question-and-answer sites rank members by their accepted answers. Generative AI can now draft, at minimal cost, the analysis, the code, and the essay that many such credentials were designed to certify (Hui et al., 2024). When the artifact behind a credential becomes cheap, the credential may certify something different from what it did, or nothing at all. When AI tools are readily accessible, the cost of producing these artifacts no longer separates the skilled ∗Olin Business School, Washington University in St. Louis. E-mail: songyao@wustl.edu. The author holds concurrent appointments as Professor of Marketing at the Olin Business School, Washington University in St. Louis, and as an Amazon Scholar. This paper describes work performed at Washington University in St. Louis and is not associated with Amazon. arXiv:2608.17111v2 [econ.GN] 14 Sep 2026 from the unskilled, and unverified skill signals already show the damage: written proposals on freelance platforms stopped predicting effort once AI could write them (Galdin and Silbert, 2025), AI-polished pitches became harder to screen (Cowgill et al., 2026), and unproctored assessment scores inflated where generative AI was accessible (Brattli et al., 2026; Serban et al., 2026). A common conclusion is that skill credentials may become worthless for analytics work (Ashton, 2026). Whether credentials backed by verified performance remain informative in the AI era, and how a platform should summarize such credentials, are open questions. Answering both requires scored outcomes before and after the AI transition. The second question is older than AI. The transition is a major shock under which that question can be tested out of sample.
We audit such a credential system end to end. Kaggle, the largest data science competition platform, has recorded 18.7 million competition entries by hundreds of thousands of participants since 2010 (B ̈onisch and Losaria, 2025), and awards medals and lifetime tiers (Expert, Master, Grandmaster) that members display as professional credentials. Every medal is won on predictions scored against withheld answers, so the credentials certify verified performance rather than artifacts.
Two features make the platform an unusually clean setting. First, two evaluation formats ran side by side: upload-competitions, which score submitted predictions and cannot tell whether the process behind them was ever validated to be generalizable, and code-competitions, which run entrants’ programs on hidden data to produce out-of-sample predictions, so a process that does not generalize is caught. The platform does not prohibit AI assistance in either format; the two formats differ in whether skipping the validation has consequences (Tadelis, 2026). Second, the platform retired the upload format for reasons that predate generative AI, which lets us watch one credential stop being issued while its stock stayed on display. Throughout, we measure one construct, the informativeness of a credential: how well it predicts subsequent hidden-test performance. Informativeness is a property of the measurement system, not of the people measured. The estimates are associational.
Whether AI changes human skill is unidentifiable here by construction, because entrants are never observed without AI access, and we make no such claim.
Four findings follow. None matches the fear that credentials are now worthless, and each informs how a platform should summarize a participant’s verified history into the credentials it displays.
First, medals are short-lived signals: in every era and format, nearly all of a medal’s predictive power sits in its first year. Second, the platform’s lifetime tiers discard up to a sixth of the information in the medals. A recency-weighted index fit on pre-AI outcomes alone explains AI-era performance better than the tiers. The index also selects entrants who perform better on average, while the tiers remain better at identifying the extreme top performers. Third, about half of the collapse of the upload-earned credential stock is institutional stranding: once the platform had phased out that format, the stock aged out under the pre-existing decay pattern. Fourth, old medals look more informative in the AI era only when read in isolation, because a stale medal proxies for its holder’s other signals, and the person-level changes we had predicted did not happen: within person, an AI-like working style predicts performance similarly in both formats, and execution failures rose with what competitions became, not who the entrants are.
Related work. This study contributes to research on reputation and certification systems on platforms. Since Spence (1973), credentials have been understood as signals that survive only while they stay correlated with what they certify. The platform literature studies how feedback and certification thresholds shape market outcomes (Tadelis, 2016; Dranove and Jin, 2010; Hui et al., 2025) and how ratings inflate or get manipulated (Filippas et al., 2022; He et al., 2022). It rarely observes the predictive content of a credential directly, and it almost never watches a credential stop being issued while its holders’ later performance keeps being scored. Both are observed here. The results complement, rather than contradict, the evidence that AI degrades unverifiable signals (Galdin and Silbert, 2025; Cowgill et al., 2026). Kaggle medals were never cheap talk, and they stayed informative through the same technology shock. The damage that did occur is a matter of design: how a platform weights aging credentials, whether it retires the format that issued those credentials, and how it displays what remains.
Method. Outcome and credentials. The outcome is one minus a team’s final percentile on the hiddentest leaderboard, so higher values mean better performance. The predictors are the participant’s medal counts prior to the competition, grouped into bands by how long ago each medal was won and which format awarded it, and entered as log(1 + m). Official tiers are reconstructed from the platform’s published deterministic thresholds and validated against observed tiers (97.7% accuracy; Appendix A5). The platform also publishes a points ranking that decays with a half-life of about a year (Appendix A2); we reconstruct it as of each competition’s start and benchmark it in section 3.2.
Discussion. 4 Implications and limitations What do hard-earned credentials still measure in the AI era, when AI tools are readily accessible?
Roughly what they always measured, for roughly as long: a medal’s predictive power still concentrates in its first year, the credentials explain about a fifth less within-competition variation than before (the tier’s R2 fell from 0.059 to 0.047), and the AI era changed how the signals were issued and read more than their informativeness. The AI era therefore enters as a test of the credentials’ informativeness rather than as a mechanism that changes it. The three design lessons below hold in both eras. What the AI transition adds is evidence that the lessons hold when the work behind the credentials changes, and that the advantage of a recency-weighted summary over the official tier survives the transition. Generative AI made polished submissions cheap to produce, but it did not make a top-rank credential cheap to achieve. Kaggle medals were never cheap talk: in either format, a medal is won on predictions verified against withheld answers. Their trajectory therefore shows what verification does not protect against: even verified signals decay fast, lose informational value when the type of evaluation that awarded them is retired, and re-weight with the signals around them.
The three platform design lessons are:
Credentials are perishable. Predictive power concentrates in the first year, so any lifetimestock display (Kaggle’s tiers, and by the same logic the badges of professional profiles on service-provider platforms) overstates stale signals by construction. The overstatement is not a corner case: in a quarter of AI-era participations by tiered entrants, every medal behind the tier is more than a year old (26%; Appendix A5). Our comparison of the official tier with the recency-weighted index puts a number on what that overstatement costs: up to a sixth (13–16%) of the available information in sample. A recent-performance summary uses no information platforms lack: an index fit before the AI era already outperforms the official tier on AI-era outcomes. Nor is it what the platform already computes: its own decaying points ranking, reconstructed as of each competition’s start, explains less AI-era variation than the frozen index and no more than the lifetime tier. The screening exercise bounds what that information recovery buys: a better broad screen, not a better instrument for the extreme top performers, where lifetime honors still win. Switching the screen from the tier to the index raises the selected entrants’ mean performance by 1.3 and 2.2 percentile points at the 10% and 25% cutoffs, respectively. The implication is a recent-performance summary displayed alongside lifetime honors, not a replacement of them. We do not observe whether buyers would use the recovered information, and the comparison is a forecasting exercise under the current rules, not a prediction of how participants would behave if the display changed.
Credential stocks are stranded by institutional exit: when a platform winds down an evaluation format, its credential stock ages out under the same decay pattern with only a moderate change in the medals’ own informativeness, and the credentials stay on display at face value. We do not claim the surviving format produces the better signal: the two formats were statistically indistinguishable in informativeness before the AI transition (0.067 against 0.065), and the platform’s shift to execution-based evaluation was already underway years before generative AI arrived. A platform that retires an evaluation format faces a stranding problem that the records themselves do not reveal.
Limitations. Several limitations bound these lessons. All estimates here are associational: they measure how well credentials predict performance, not why. We also do not observe the demand side: no buyer or employer is seen responding to a credential, so the paper establishes what the credentials predict rather than how they are used, and that demand response is the natural next step. One further gap is an extension rather than a flaw: Kaggle is one platform, chosen because it ran two evaluation formats side by side and phased one out, which is what makes the audit possible. The three lessons are properties of how credentials are aggregated and displayed, so they should carry to other platforms with publicly displayed lifetime credentials that buyers use to screen providers. The finding that verified signals kept their informativeness through the AI transition is narrower. Cheap talk, such as a polished r ́esum ́e, loses its value as a signal once AI can produce the polish for anyone.
A hard-earned credential is scored against withheld answers, which polish cannot move. The same analyses can be repeated on any such platform if the data become available. Two limitations concern the AI-era sample itself.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What gaps exist between benchmark performance and real deployment outcomes?- How much does medal age matter when predicting Kaggle performance?
- Do medals from retired competition formats lose predictive power faster than others?
- Should platforms downweight or relabel credentials after retiring the format that issued them?
- Does this stranding problem appear when other evaluation platforms retire their formats?
- What makes a credential robust when the tools for earning it change?
- What counts as evidence that a credential still certifies after GenAI?
- What makes counterfeiting social warrant different from counterfeiting factual claims?
- How do verification labels themselves become part of the misinformation problem?
- Why do intellectual products gain false authority from AI-generated form?
- Can AI output be verified without understanding the reasoning behind it?
- What does it mean that AI knowledge is structurally hearsay?
- Why is AI output fundamentally unverifiable against underlying reality?
- How does AI knowledge become structurally different from written sources?
- What happens when AI generates content faster than humans can verify it?
- Can users interrogate AI outputs without verifying every single claim?
- What verification methods work for knowledge without stable referents?
- How can we verify outputs from systems that generate without grounding?
- Can AI evaluation tools solve the verification problem they help create?
- Why does AI generation outpace verification across the research lifecycle?