Can live benchmarks prevent data contamination in prediction tasks?
How can prediction benchmarks stay contamination-free when future outcomes aren't yet known? FutureX tests whether continuous real-time updates eliminate training data leakage.
Future prediction is a hard agent task — analytical thinking, information gathering, decision under uncertainty — and until FutureX there was no large-scale benchmark for it, largely because real-time updates and timely answer-retrieval are hard to operate. FutureX's design choice is the keeper: it is a live benchmark that continuously collects questions from 195 trusted sites, gathers model predictions at each event's start date, and automatically checks actual outcomes. Being live is not a convenience — it is the contamination defense: a benchmark whose answers don't exist yet cannot leak into training data.
The capability finding across 25 models is equally clean: strong base models (e.g., DouBao-Seed1.6) handle straightforward questions, but hard open-ended prediction requires built-in search and reasoning, with deep-research and Think&Search agents (Grok-4, GPT-o4-mini) leading on the hardest tasks. Forecasting is therefore an agentic capability, not a base-model one.
This pairs directly with Batch 1's evaluation thread. Since Do automated benchmarks hide what frontier AI systems can really do?, FutureX is a concrete open-world instrument whose live-updating mechanism operationalizes the contamination-free, real-task ideal; and it complements Can frontier exams really measure cutting-edge AI capability? — where HLE restores discrimination on static knowledge, FutureX restores it on dynamic prediction. It also grounds Can LLMs actually forecast time series better than we think?: the gain comes from the search-and-reason workflow.
Inquiring lines that read this note 8
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do capability benchmark scores systematically misrepresent true model abilities?- Why are post-cutoff test sets essential for evaluating genuine forecasting ability?
- What real-world tasks most clearly expose gaps between benchmark performance and actual capability?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Why do benchmarks become saturated so quickly after initial launch?
- How does contamination protection by time differ from protection by scarcity?
- How can hidden test partitions detect constant predictions that generalize?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
FutureX is a live, contamination-free instance of the open-world evaluation ideal
-
Can frontier exams really measure cutting-edge AI capability?
Popular benchmarks like MMLU saturate quickly, hiding real capability differences. Can expert-designed closed-ended exams like Humanity's Last Exam discriminate at the frontier, and what would high scores actually tell us about AI systems?
complementary: static frontier exams vs dynamic prediction
-
Can LLMs actually forecast time series better than we think?
Explores whether language models possess stronger forecasting ability than current benchmarks suggest, and what role workflow design plays in revealing or hiding that capability.
both find forecasting gains come from agentic workflow not base model
-
Can scarcity of solutions protect benchmarks from data contamination?
ExploitGym lacks ground-truth exploits for many tasks, which might prevent models from memorizing solutions during training. But does difficulty-based protection actually hold up, or does it degrade once solutions become public?
a second contamination defense, by scarcity of solutions instead of liveness; it may erode if solutions are later published, and it costs the ability to tell a failure from an impossible task
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
- Task Contamination: Language Models May Not Be Few-Shot Anymore
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- Can Large Reasoning Models Self-Train?
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
Original note title
future-prediction benchmarks must be live and contamination-free and open-ended forecasting requires search-and-reasoning agents not base models