The Future of Facts: Tracing the Factual Generation-Verification Gap
Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generationverification gap (GV-gap) underlies many recent advances in self-improvement and reasoning, but its dynamics on factual knowledge specifically remain poorly understood. We focus on the training mechanisms underlying factual GV-gaps, distinguishing them from their computational and aesthetic counterparts. We trace generation and verification capabilities through three training phases (acquisition, continual learning, and updating) across four open-source model families at two scales each. Three findings recur across models: (i) verification is consistently learned before generation; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a multi-verse state, simultaneously verifying both old and new answers as correct. Natural experiments on frontier models reproduce these dynamics at scale and reveal residual verification biases on well-covered facts. †
Introduction. Several researchers and practitioners have observed that language models (LMs) are often more accurate at verifying outputs than at generating them [1, 2, 3, 4]. For example, given a factual triplet such as “{Paris, IsCapitalOf, France}”, an LM may be more accurate when asked the verification query “Is Paris the capital of France?” than the generative one “What is the capital of France?” While this “generation-verification gap” (GV-gap) underlies a wide range of recent innovations in self-improvement and reasoning [e.g., 5, 6, 7, 8, 9] and has attracted growing theoretical attention [4, 10, 11], empirical accounts of how the training process shapes the gap on factual knowledge — how it emerges, persists, and changes across the life cycle of a fact — remain limited and partially conflicted.
This blindspot is increasingly consequential. LMs are rapidly becoming the default interface to factual knowledge, with billions of users querying systems like ChatGPT [12] or encountering “AI summaries” beneath every search bar [13, 14]. As LMs come to mediate what people read and verify, structural asymmetries in what they produce, affirm, and deny shape the information environment itself. These concerns are sharpened by ongoing worries about AI-generated content becoming indistinguishable from authentic media [15, 16] and by models’ tendency to “hallucinate” confident but false outputs [17, 18, 19]. When users querying different LMs receive conflicting factual responses, the resulting fragmentation feeds broader concerns about an emerging “post-truth” world [20, 21, 22].
Studying how training shapes the factual GV-gap is empirically difficult. Training recipes for frontier models are largely opaque, which is true even for many “open-source” models where data composition is rarely fully disclosed. Even when training data is released, its sheer scale makes it near Existing work touches the factual GV-gap from several angles but rarely studies it as a trainingmechanism phenomenon. A large literature examines the model mechanisms of factuality — how facts are learned [26, 31, 32, 33] and retrieved [34, 35, 36, 37], how knowledge can be “edited” [36, 38, 39], and how continual learning can induce (catastrophic) forgetting [40, 41, 42] — but generally treats factual knowledge as a single capability rather than separating generation from verification. Complementary lines of research formalize the GV-gap and quantify its magnitude on popular benchmarks [2], or leverage it for self-improvement and test-time scaling [7, 8, 27]; these illuminate the gap’s potential but, being tied to opaque training data, offer little insight into its origins. Theoretical accounts [4, 10] provide useful framing but are not grounded in life-cycle dynamics of factual data.
Our work contributes a controlled study of the training mechanisms underlying the factual GV-gap, alongside a taxonomy that situates factual GV-gaps relative to their computational and aesthetic counterparts. Using synthetic facts in domains relevant to public discourse, we fine-tune four opensource model families across two scales each and trace both capabilities through three training phases: acquisition, continual learning, and updating. Three findings recur across models and scales: (i) verification is consistently learned before generation, exposing a window in which models reliably verify facts they cannot yet produce; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a “multi-verse” state, simultaneously verifying both the old and the new answer as correct. Natural experiments on flagship frontier models, exploiting variation in real-world data coverage across topics and time, reproduce the same dynamics at scale and reveal residual verification biases on well-covered facts.
Our findings situate the factual GV-gap within the model training process, complementing research into model mechanisms. As LMs increasingly produce the content that trains the next generation of models, these asymmetries risk compounding through the data they leave behind. Addressing this calls for better choices in data curation, training curricula, and evaluation. To support such work, we open-source our full experimental setup as a shared testbed for studying the factual GV-gap.
Method. The term generation-verification gap is used across the LM literature to describe several distinct phenomena that share a common shape: models are better at verifying outputs than producing them. We distinguish at least three classes that differ in what is being verified, how cleanly outcomes can be measured, and how directly capabilities can be traced to specific training data.
Factual. The distinction between factual recall (generation) and recognition (verification), e.g., remembering a name versus picking it from a list, has long been studied in cognitive science [43], where dual-process models posit related but partially distinct mechanisms for the two [44, 45, 46, 47]. A parallel asymmetry appears in statistical learning theory: learning a discriminative decision boundary is generally easier than learning the full joint distribution over inputs and outputs [48, 49]. In LMs, the asymmetry is plausibly amplified by the structural difference between the two operations: verification typically reduces to a decision over a small token space (e.g., a binary True/False), while autoregressive generation requires sampling a sequence from the joint distribution over the full vocabulary, with each step compounding the difficulty [50]. Crucially, both capabilities can be traced to specific training data points and their strength objectively measured, properties that make factual GV-gaps a clean target for empirical study.
Computational. The computational GV-gap is rooted in the classical P–vs–NP distinction: for many problems, generating a solution is provably harder than checking one [51, 52, 53]. Outcomes are cleanly measured against ground truth, but tracing the gap back to specific training data is difficult. As Ruis et al. [23] note, LMs’ procedural reasoning appears to draw on a large, scattered set of training examples rather than a few identifiable sources.
Aesthetic. Humans can recognize the brilliance of Shakespeare’s prose or the beauty of the Sistine Chapel without being able to explain, let alone reproduce, what makes them great. Similarly, LMs have shown capable of preferring outputs from stronger models over their own [30], while being unable to match their quality. Unlike the factual and computational cases, aesthetic outputs resist objective measurement [54], and the underlying capability presumably draws on training data even more diffuse than that supporting procedural reasoning.
We focus on factual GV-gaps for their direct relevance to how LMs are used as knowledge interfaces and the methodological tractability of factual data. Returning to the triplet (Paris, IsCapitalOf, France): a generative failure occurs when an LM produces an incorrect answer (e.g., “London”) to ”What is the capital of France?”, a phenomenon commonly called hallucination [17]. A verification failure occurs when the model classifies the correct statement “The capital of France is Paris” as wrong, or accepts the incorrect statement “The capital of France is London” as right — the latter associated with sycophancy [55]. Two complementary lenses on the gap follow naturally from this setup. The user-facing search-engine lens, our primary focus, treats the LM as a knowledge base queried by users. The model-facing self-improvement lens treats the LM’s own generations as candidate outputs to be verified.
Search-engine lens. This lens models a user querying an LM m for facts drawn from a dataset of triplets (x, r, y∗) ∼D, with ̃y ∼ ̃Y(x, r) denoting a plausible incorrect candidate for the same (x, r). We define generative and verification utilities as:
Self-improvement lens. Here, the candidate y is sampled from the model itself rather than provided by a user, so the relevant error rate is αm = 1 −UG(m, D). For a single sample y ∼m(x, r) this yields the self-consistency utility Drawing multiple samples per query recovers the Best–of–N rejection-sampling regime, where ∆ measures how much the model’s verifier improves over its own generator.
Discussion. Implications. Most new content on the web will soon be produced or mediated by LMs [75, 76], meaning the data that train tomorrow’s models increasingly reflect the asymmetries of today’s. A factual GV-gap that quietly favors verification over generation, and lingers across updates, is therefore not a local property of any one model but a global dynamic shaping the shared information substrate [77]. Unlike outright model collapse on recursive data [78], this is a subtler structural asymmetry in what models verify as true. Our controlled experiments show the gap converges with sufficient exposure and that frontier models perform well on well-covered facts. Current shortcomings are therefore unlikely to be fundamental to transformers, but signal progress is achievable through better data curation and training strategies.
This optimism has limits at the small end. Parametric memory is fundamentally bounded [79], and small models can thus not be expected to store all factual knowledge their users may ask about. The desired behavior in such cases is abstention rather than confabulation, a strategy not generally rewarded by most benchmarks [80]. Indeed, we find that even when small models like GPT-5.4 Nano do abstain on generation queries, they do not extend this behavior to verification (App. H.4.5).
Rethinking how we measure factuality. Our findings suggest that when it comes to factuality, verification and generation should be treated as distinct capabilities with different learning dynamics, rather than two facets of the same underlying competence. Yet current evaluation practice rarely measures them separately, and the data-exposure thresholds at which each emerges are not visible from training loss alone (Section 4.1). A better understanding of this gap calls for testbeds and datasets that track factual capabilities not only at a frozen checkpoint on a frozen set of facts, but as both models and facts evolve. It also calls for mechanistic model accounts of the gap, complementing our training-mechanism framing; our setup offers a controlled probe for such mechanistic studies.
Conclusion. Language models are replacing search engines as our default interface to factual knowledge, but they do not treat knowledge uniformly. This work focused on the factual GV-gap, whose origins trace to specific data and whose strength can be objectively measured. Across four open model families and a controlled training life cycle of synthetic facts, we find that verification is learned before generation, survives continual learning more robustly, and may enable a “multi-verse” where superseded answers remain verified as true. Naturalistic experiments on deployed flagship models reproduce these regimes, and added reasoning effort does not appear to close the gap. We do not believe the gap is a fundamental limit of transformers, but rather a limitation of how we curate data and train models. As we offload more factual cognition to LMs, more of what they read tomorrow is what they wrote today — a recipe for a future of proliferating factual disputes and a quiet drift in what counts as usable knowledge. By open-sourcing our setup for reliable fact injection we hope to give the community a shared instrument for studying the factual GV-gap.
Limitations. We focused on single-hop facts as a clean testbed for isolating GV-gap dynamics; multi-hop facts entangle factual recall and recognition with multi-step reasoning [57], and we leave their study to future work. Our facts are also presented as short paraphrased sentences explicitly containing each triplet, cleaner and denser than naturally occurring training data. We conjecture this makes our setup a conservative lower bound on the dynamics we report. Our setup also injects facts during post-training rather than pretraining. We view it as representative of the increasingly common “mid-training” phase [63, 81], and our naturalistic results suggest similar dynamics surface at the frontier scale, although a direct test at pretraining scale remains open. Recent frontier models also started to recognize evaluation contexts [82], a confounder we also observed (Fig. H.6) that complicates interpretation for the most capable models. Finally, we did not evaluate mitigations such as retrieval-augmented generation (RAG) [83] or self-improvement methods like Best-of-N. RAG can sidestep some of the dynamics we identify but adds retrieval and serving overhead, increasingly relies on LM-generated source material, and still requires verification capabilities to select among retrieved candidates.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why does AI verification capability persistently exceed generation capability?- Why does verification take longer than generation across research workflows?
- How much does model verification capability exceed generation capability?
- How does the generation-verification gap change as models scale up?
- Why does research artifact generation outpace verification while facts show the opposite pattern?
- Why does the right structural prior matter more than raw model capacity?
- Why does tool use decouple factual capacity from model parameter count?
- How do model size and document diversity interact in SDF override success?
- Why does in-weight memorization fail compared to tool-based fact access?
- Why does attending to own latents work better than bolted-on external memory stores?
- How does in-weight memorization scale with model parameter count?
- What is the theoretical capacity limit before memorization saturates?
- Is forgetting in language models reversible or permanent knowledge loss?
- How do newly learned facts become accessible after gradient updates?
- What makes factual memorization less efficient than tool-based retrieval?
- Can in-weight memorization scale beyond model parameter count limits?