When an AI lab reports a benchmark score, is it measuring the model or the setup it ran in?
How much do internal benchmarks differ from independent model evaluations?
This explores how far the benchmark scores AI labs report for their own models can drift from what independent evaluators measure, and why the two might not match.
This explores how far the scores labs report for their own models can drift from what independent evaluators find. The corpus doesn't have a study that directly compares lab-reported numbers with third-party re-runs, so it can't give you a size for the gap. What it does have is an explanation of where the gap comes from. A benchmark score belongs to a model running inside a particular setup, not to the model alone. When the setup changes, the score changes too, even though the model hasn't.
The clearest evidence comes from work on "harnesses," the scaffolding of tools, memory and retry logic that a model runs inside. Optimizing only that scaffolding around frozen models lifted Terminal-Bench 2.1 scores across several models, including a 5.4-point gain for DeepSeek-V4 Flash, with no change to the weights Can execution harnesses lift model performance without retuning weights?. A small 9B model trained to patch harnesses raised a frozen agent's performance by 9.3 points Does training editors on real outcomes beat prompting larger models?. Another system gives models a layered external memory and reports gains on ARC-AGI-3, but its authors haven't tested which parts drive the improvement Can external state caches let models solve harder problems?. So if a lab tests its model in a carefully tuned in-house setup and an independent group uses a generic one, a gap of several points can appear without anyone being dishonest.
That is why provenance matters. Benchmark Radar, a living catalog of evaluations, keeps each score tied to its source and citation. Its argument is that a score stripped of its origin can't support a fair comparison between models Can benchmark scores be trusted without knowing their origin?. You might expect richer, interactive agent benchmarks to fix this. They don't: comparability and reproducibility problems reappear at the level of whole agent runs, now with more settings that can quietly differ Do interactive evaluations actually solve the benchmark comparison problem?.
There's a second, less obvious layer. Two models with the same score can be built very differently inside. One may hold together under small changes to the input or unfamiliar data while the other breaks, and standard metrics don't show which is which Can models be smart without organized internal structure?. Where an independent evaluator uses slightly different prompts or data, it is effectively running a robustness test the original benchmark never ran. Results on instruction-following show how sharply this can play out. Reasoning models hold up to roughly 150 simultaneous instructions and then fail steeply, so a test that sits just below or just above that point will tell very different stories about the same model How does instruction density affect model performance?.
The takeaway is that "internal vs. independent" is often less about who runs the test than about what the test includes: which harness, which prompts, how hard, and whether the original setup is recorded. Before trusting any reported number, ask what the model was run inside. One vocabulary warning: in test-time-scaling research, "internal" means something unrelated, namely reasoning built into the model through training, as opposed to search and checking added at inference time How do internal and external test-time scaling compare?.
Sources 8 notes
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Prime Agent organizes persistent state in four levels (weights, context, persistent REPL plus subagents, disk-backed history) to let models read and write addressable state beyond their instruction stream. The approach isolates harness failures from model failures and reported gains on ARC-AGI-3, though specific components remain unablated.
Benchmark Radar catalogs AI evaluations while keeping source identities and citations attached, enabling readers to trace scores back to their original settings. The system demonstrates that scores without provenance cannot reliably support model comparisons.
Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.
Show all 8 sources
Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.
IFScale benchmark shows three degradation patterns: linear (small models), exponential (mid-range), and threshold decay (reasoning models maintain ~150 instructions then fail steeply). Even best models reach only 68% accuracy at maximum density.
Research shows test-time scaling methods split into internal (training models for autonomous reasoning) and external (inference-time search and verification). They complement rather than compete; internal builds capability while external extracts performance from existing capability.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Interactive Evaluation Requires a Design Science
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Sharpening Tax in Post-Training
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- How Many Instructions Can LLMs Follow at Once?
- Prime Agent: A Self-Improving RLM Harness