When a contest accepts uploaded predictions instead of code, can its winners be tuned to the test rather than the task?
Why did upload competitions fail to catch non-generalizable predictions that code competitions caught?
This explores why machine learning contests where entrants upload finished predictions let overfit or leaked results through, while contests that require entrants to submit runnable code exposed those results as non-generalizable.
This explores why contests that accept uploaded prediction files let fragile, non-generalizable results through, while contests that run participants' code on data they never see caught them. Up front: the collection has no note that studies this comparison directly. No retrieved source discusses upload-versus-code competition formats, so what follows comes from adjacent material and explains the likely mechanism. It is not a documented finding.
The closest conceptual match is the idea that gaming a score is one failure that shows up in different places. Does reward hacking always stem from the same failure? argues that reward hacking happens wherever something is optimized against a signal that only partly represents the real task. It doesn't matter whether the optimizing happens in training, in picking outputs, or in revising prompts. An upload leaderboard is one of those signals. If entrants can see the test inputs, or can submit many times and watch the score move, they can tune toward that specific test set instead of the underlying problem. The leaderboard rewards fitting the exam. A code competition shrinks that gap: your method has to run on data it has never seen, so the score measures the method rather than the predictions.
The collection's material on evaluation that holds up offers a useful contrast. Can crowdsourced votes reliably rank language models? credits Chatbot Arena's rankings to a steady supply of fresh, diverse, discriminating questions that no model can prepare for in advance. That is the same protective property a hidden code-competition test set has. A related idea appears in Can past performance predict when a model will be right?: how well a model performs on one case tells you little about how reliable it is. Reliability shows up across many outcomes over time. An uploaded prediction file is a single snapshot, while code that gets re-run on new data adds to a track record.
The less obvious lesson is that the format of an evaluation is part of what it measures. Changing how answers are submitted, without changing the task, changes which kinds of cheating or overfitting can survive. If you want primary sources on competition design itself, such as leaderboard probing, test-set leakage or Kaggle's move to code competitions, they are a gap in this collection.
Sources 3 notes
Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.
Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.
XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Deep Think with Confidence