INQUIRING LINE

Do AI models get better at spotting right answers faster than they get better at producing them as they grow?

How does the generation-verification gap change as models scale up?

This explores whether a model's ability to check answers and its ability to produce them grow apart or move closer together as models get bigger, and whether scale is even what drives that gap.


This explores whether bigger models get better at checking answers faster than they get better at producing them, or the other way around. The collection doesn't contain a clean curve plotting the generation-verification gap against model size, so a direct answer isn't available. What it does have is more interesting: several findings suggest that parameter count isn't the main thing controlling the gap.

The closest evidence is that verification is learned before generation. Across four model families at several sizes, models learn to recognize a correct fact earlier in training than they learn to produce it. That recognition also holds up better when the model is later updated, so much so that an updated model can accept both the old fact and the new one as correct Why do models verify facts better than they generate them?. Because this pattern appears at every scale studied, the gap looks like a basic property of how these models learn rather than something growth erases. Saying yes or no to a candidate is an easier learning problem than producing a whole correct sequence. One surprising consequence: a better checker is not always a better guardian, because a strong verifier can be confidently wrong about which version of a fact is current.

That same gap is why models can't simply improve themselves. Pure self-improvement stalls because a model's judgment of its own work isn't reliable enough to learn from. The methods that do work quietly bring in an outside reference point, such as an earlier model version, a separate judge, user corrections, or tool feedback Can models reliably improve themselves without external feedback?. A related finding shows that LLMs are good at proposing candidates but poor at estimating how valuable or uncertain those candidates are. In scientific discovery, they needed a statistical model fitted to real experimental data to fill that role Can language models reliably judge their own candidate quality?.

Here is the twist: verification may scale along its own axis, separate from model size. Verification accuracy improves with finer-grained scores, repeated evaluation, and breaking judgments into separate criteria, all applied at inference time without any retraining. On this view, weak verifiers are under-resourced rather than fundamentally limited Can verification accuracy scale without training models?. Small verifiers that reason before giving a verdict show how far this can go: a 1.5B-parameter process reward model beats GPT-4o at judging reasoning steps, and another reaches state-of-the-art results with 1% of the usual training labels Can generative reasoning beat discriminative models with less training data?. Verification can also run alongside generation without slowing it down Can verifiers monitor reasoning without slowing generation down?.

The same pattern shows up on the generation side. Extra compute at inference time can stand in for a larger model on hard problems Can inference compute replace scaling up model size?. A 3B model with a well-designed post-training pipeline matches frontier reasoning scores, but only on tasks where answers can be checked automatically Can small models match frontier reasoning without massive scale?. So verification is what lets small models catch up, rather than something scale has to supply. Scale also doesn't help evenly. In one study, the ability to benefit from improvements to an agent's harness (the scaffolding around the model) peaked at mid-sized models: weak models didn't use the harness, and strong ones didn't follow its instructions faithfully Do stronger models always evolve harnesses better?. The takeaway is to stop asking how scale changes the gap and start asking where verification capacity comes from, since much of it can be added without making the model any bigger.


Sources 9 notes

Why do models verify facts better than they generate them?

Across four model families and scales, verification accuracy develops earlier in training than generation, remains more robust to continual learning, and can leave updated models accepting both old and new facts as correct simultaneously. This asymmetry reflects different learning difficulties: verification requires binary decisions while generation requires sampling full sequences.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Can language models reliably judge their own candidate quality?

LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.

Can verification accuracy scale without training models?

Research shows verification accuracy improves independently via score granularity, repeated evaluation, and criteria decomposition—all deployable at inference without retraining. This reframes weak verifiers as under-scaled rather than fundamentally limited.

Can generative reasoning beat discriminative models with less training data?

GenPRM and ThinkPRM reframe process supervision as generative tasks with CoT reasoning before judgment, achieving superior performance on far fewer labels. A 1.5B GenPRM beats GPT-4o; ThinkPRM uses only 1% of PRM800K labels to surpass full-dataset discriminative verifiers.

Show all 9 sources
Can verifiers monitor reasoning without slowing generation down?

Decoupling verification from generation lets verifiers run alongside a single trace, forking to extract verifiable state and intervening only on violations. On correct runs the latency penalty is near-zero; interwhen matches or beats CoT across benchmarks at similar token budgets.

Can inference compute replace scaling up model size?

Snell et al. (2024) showed that inference-time compute trades off against model parameter scaling, especially on difficult prompts. This reveals pretraining and inference compute are not independent resources.

Can small models match frontier reasoning without massive scale?

A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.