SYNTHESIS NOTE
Topics›Evaluations›this note

Can live benchmarks prevent data contamination in prediction tasks?

How can prediction benchmarks stay contamination-free when future outcomes aren't yet known? FutureX tests whether continuous real-time updates eliminate training data leakage.

Synthesis note · 2026-06-03 · sourced from Evaluations
How does test-time scaling work for individual research agents?

Future prediction is a hard agent task — analytical thinking, information gathering, decision under uncertainty — and until FutureX there was no large-scale benchmark for it, largely because real-time updates and timely answer-retrieval are hard to operate. FutureX's design choice is the keeper: it is a live benchmark that continuously collects questions from 195 trusted sites, gathers model predictions at each event's start date, and automatically checks actual outcomes. Being live is not a convenience — it is the contamination defense: a benchmark whose answers don't exist yet cannot leak into training data.

The capability finding across 25 models is equally clean: strong base models (e.g., DouBao-Seed1.6) handle straightforward questions, but hard open-ended prediction requires built-in search and reasoning, with deep-research and Think&Search agents (Grok-4, GPT-o4-mini) leading on the hardest tasks. Forecasting is therefore an agentic capability, not a base-model one.

This pairs directly with Batch 1's evaluation thread. Since Do automated benchmarks hide what frontier AI systems can really do?, FutureX is a concrete open-world instrument whose live-updating mechanism operationalizes the contamination-free, real-task ideal; and it complements Can frontier exams really measure cutting-edge AI capability? — where HLE restores discrimination on static knowledge, FutureX restores it on dynamic prediction. It also grounds Can LLMs actually forecast time series better than we think?: the gain comes from the search-and-reason workflow.

Inquiring lines that read this note 8

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do capability benchmark scores systematically misrepresent true model abilities? How does evaluation scope and dimensionality affect what we measure? Why do people disclose to AI systems despite their artificial nature?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 112 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

future-prediction benchmarks must be live and contamination-free and open-ended forecasting requires search-and-reasoning agents not base models