Does a speedup on a small AI model actually predict what happens when you scale it up to frontier size?
Can frontier-scale results be predicted from small-scale benchmark speedups?
This explores whether a gain measured on small models or cheap benchmarks (a training trick, a speedup, a better score) tells you what will happen when the same idea runs at frontier scale. The corpus doesn't answer that directly, but it has a lot on why small-scale signals transfer in some cases and break in others.
This explores whether a win measured cheaply, on small models or quick benchmarks, can tell you what happens at frontier scale. The direct answer first: this collection has no paper that tests scaling-law extrapolation or speedrun-to-frontier transfer head-on. What it does have is several cases that show when small-scale signals carry over and when they mislead. The short version is that the shape of the curve can change as models grow, so a straight-line extrapolation is risky.
The clearest warning comes from instruction-following research. When you pile more and more instructions into a prompt, small models degrade linearly, mid-sized models degrade exponentially, and reasoning models hold steady until around 150 instructions and then fall off a cliff How does instruction density affect model performance?. If you only measured small models, you'd predict a gentle slope and miss the cliff entirely. The failure pattern itself depends on the class of model, so a small-scale result can predict the wrong kind of behavior, not just the wrong number. The same lesson shows up between reasoning and non-reasoning models: extra inference compute doesn't close the gap, because training sets up a protocol that makes extra tokens useful Can non-reasoning models catch up with more compute?. Gains that look interchangeable at one scale may not be interchangeable at another.
In the other direction, scale is a messier axis than it looks. A 3B model with a well-designed post-training pipeline matches much larger systems on math and code Can small models match frontier reasoning without massive scale?. Smaller models given more inference compute can match bigger ones on hard prompts Can inference compute replace scaling up model size?. Routing queries among several 7B models has beaten frontier models outright Can routing beat building one better model?. If parameter count can be traded for pipeline design, inference budget, or routing, then "small-scale" and "frontier-scale" aren't two points on one line, and predicting from one to the other means knowing which lever produced the gain. One caveat matters here: these wins are concentrated on tasks with checkable answers, where clean reward signals exist.
Some improvements do seem to travel well. Improvements to the scaffolding around a frozen model, the "harness," carried over to newer models without changes Can execution harnesses lift model performance without retuning weights?. AIDE2's gains held up on held-out benchmarks, including weather forecasting, which sat outside its training distribution Do AIDE2's improvements transfer to unseen tasks?. The pattern suggests that changes to the system around the model transfer more reliably than changes measured inside one model at one size. That's also why a small model trained on real patch outcomes beat prompted frontier models at harness editing: it checked its work against actual results instead of guessing Does training editors on real outcomes beat prompting larger models?.
The biggest catch is that the benchmark itself may be the weak link. Automated benchmarks both overstate and understate what frontier systems can do, because they favor tasks that are tightly specified and easy to grade Do automated benchmarks hide what frontier AI systems can really do?. Hard expert exams only separate models for a while before they saturate Can frontier exams really measure cutting-edge AI capability?. So a small-scale speedup on a benchmark is a prediction made through two lenses at once: from small to large, and from benchmark to reality. The collection gives you more reasons to distrust the second lens than the first.
Sources 10 notes
IFScale benchmark shows three degradation patterns: linear (small models), exponential (mid-range), and threshold decay (reasoning models maintain ~150 instructions then fail steeply). Even best models reach only 68% accuracy at maximum density.
Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.
A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.
Snell et al. (2024) showed that inference-time compute trades off against model parameter scaling, especially on difficult prompts. This reveals pretraining and inference compute are not independent resources.
Avengers-Pro achieves 7% higher accuracy than GPT-5-medium by routing queries to optimal models per semantic cluster, or matches its performance at 27% lower cost. Ten 7B models with routing previously surpassed GPT-4.1 and 4.5, suggesting selection is a stronger lever than scaling.
Show all 10 sources
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
Humanity's Last Exam uses 3,000 expert-designed questions to expose capability gaps where MMLU saturates, showing real discrimination—but expert exam performance wouldn't indicate autonomous research or open-world problem-solving that matters for deployment.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Reasoning Models Can Be Effective Without Thinking
- Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
- Agents' Last Exam
- Open-World Evaluations for Measuring Frontier AI Capabilities
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Sharpening Tax in Post-Training