Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence

Paper · arXiv 2510.16657 · Published October 18, 2025
Frontier AI Risk & RSI

Synthetic data has been increasingly used to train frontier generative models. However, recent studies raise key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model performance, a phenomenon often coined model collapse. In this paper, we investigate ways to modify the synthetic retraining process to avoid model collapse, and even possibly help reverse the trend from collapse to improvement. Our key finding is that by injecting information through an external synthetic data verifier, whether a human or a better model, synthetic retraining will not cause model collapse. Specifically, we situate our theoretical analysis in the fundamental linear regression setting, showing that verifier-guided retraining can yield near-term improvements, but ultimately drives the parameter estimate to the verifier’s “knowledge center” in the long run. Our theory further predicts that, unless the verifier is perfectly reliable, these early gains will plateau and may even reverse. Indeed, our experiments across linear regression, Variational Autoencoders (VAEs) trained on MNIST, and fining-tuning SmolLM2-135M on the XSUM task confirm these theoretical insights.

Introduction. The use of synthetic data has gained significant traction due to its ability to reduce data collection costs and enhance privacy protection, with applications in computer vision (Wood et al., 2021), healthcare (Azizi et al., 2021; Santangelo et al., 2025), finance (Potluru et al., 2023) and recently large language models (LLMs) (Chen et al., 2024). A growing body of work has demonstrated that training with synthetic data can improve model performance in various applications such as image recognition (He et al., 2023; Tremblay et al., 2018), image generation (Doersch & Zisserman, 2019; Shrivastava et al., 2017; Tian et al., 2023) and language generation (Gunasekar et al., 2023; Guo et al., 2024; Zelikman et al., 2022). However, recent studies caution that recursively training models on synthetic data alone can lead to a degradation of quality, a phenomenon often coined model collapse (Shumailov et al., 2024; Dohmatob et al., 2024a, 2025, 2024b; Alemohammad et al., 2024; Gerstgrasser et al., 2024). This contrast between empirical success and pessimistic research findings gives rise to a natural question about what might have caused the discrepancy.

In response to the above question, an important practical observation is that synthetic data are rarely used in raw form in the aforementioned applications. Instead, practitioners often apply filtering steps to remove low-quality synthetic samples before retraining. For example, in natural language generation, synthetic sentences are often screened using grammar checkers or LLM-as-a-judge pipelines (Zheng et al., 2023; Gu et al., 2024); in computer vision, synthetic images can be filtered using pretrained discriminators or human annotation (He et al., 2023); in recommendation and preference learning, synthetic feedback is often validated against external heuristics or known user signals (Tu et al., 2025; Iskander et al., 2024; Lupidi et al., 2024; Lampis et al., 2023; Zhang et al., 2024). Common across all these approaches is the use of a knowledgeable “discriminator” (henceforth the verifier ) – whether a machine or human – that evaluates and filters out low-quality candidate synthetic samples (i.e., those not passing the discriminator’s screening). This observation naturally raises the following research question:

Does verifier-based filtering of synthetic data contribute to the observed empirical success of model improvements, and does it prevent model collapse in the long run?

There has been a growing body of work studying the mechanisms of model collapse, including theoretical analyzes that are often examined in the context of classic parameter estimation problems (Dohmatob et al., 2024a, 2025; Gerstgrasser et al., 2024; Dey & Donoho, 2024; Xu et al., 2025; Suresh et al., 2025). However, these works have all assumed the use of synthetic data without filtering. Only few recent works start to examine how filtered synthetic data affects the performance of iterative retraining, but under idealized assumptions such as access to a perfectly correct verifier (Amin et al., 2025) or highly structured errors in synthetic data (i.e., i.i.d. noise added to binary labels (Feng et al., 2025)). A realistic and instructive framework for analyzing synthetic retraining of generative models remains still poorly understood. In this paper, we further push the boundary of this important agenda, and examine iterative retraining under imperfect verifiers that filter out low-quality synthetic data based on their (possibly biased) knowledge. We refer to this process as verifier-based synthetic retraining for convenience. Specifically, we seek principled understandings about the empirically observed short-term successes of verifier-based synthetic retraining and analysis about its long-term convergence under iterative retraining.

Our contributions. We start from theoretical investigations and analyze verifier-based synthetic retraining with verified synthetic data on the foundational linear regression model—a canonical setting that has become central to the study of model collapse (e.g., (Dohmatob et al., 2024a, 2025; Gerstgrasser et al., 2024; Zhu et al., 2025; Garg et al., 2025)).2 We then verify our theoretical insights through thorough empirical studies in real-world generative settings. Our main contributions are summarized as follows:

• Does verified synthetic data improve retraining and, if so, under what conditions? We show that it indeed can, provided the right conditions are met. Through a new form of bias-variance trade-off under data filtering, we characterize the regimes in which verifier-based synthetic retraining leads to strict model improvement, rather than degradation, in the short term (Theorem 3.1). The conditions we identify highlight the mixed effect of synthetic sample size, the verifier’s bias and selectivity during filtering, yielding practical insights regarding when verification of synthetic data is beneficial.

Related work. Understanding and mitigating model collapse. Recent work shows that heavy reliance on synthetic data in iterative training can cause model collapse—the degradation of performance when a model is repeatedly retrained on its own synthetic outputs (possibly mixed with real data).3 Empirical evidence supports this phenomenon: Shumailov et al. (2024) show that recursive training on unfiltered synthetic data induces distribution shift and mode collapse, while Dohmatob et al. (2025) find that even small synthetic proportions can harm performance. In linear settings, Dohmatob et al. (2024a) analyze collapse mechanisms explicitly, and Dohmatob et al. (2024b) connect degradation to altered neural scaling laws.

To mitigate collapse, prior work broadly explores three strategies. First, accumulating data or gradually increasing the synthetic dataset size across iterations can suppress noise and bound errors (Gerstgrasser et al., 2024; Dey & Donoho, 2024; Xu et al., 2025; Kazdan et al., 2025; Barzilai & Shamir, 2025). Second, mixing synthetic data with real data stabilizes retraining (Bertrand et al., 2024; Fu et al., 2024, 2025), as performance progressively degrades without sufficient fresh real data (Alemohammad et al., 2024). Recent studies have even derived optimal mixing ratios to maximize this stabilizing effect (He et al., 2025; Garg et al., 2025). Finally, algorithmic interventions, such as the token re-sampling procedures proposed by Zhu et al. (2025), offer alternative pathways to avoid collapse.

Unlike prior work that relies on unfiltered synthetic data, our framework incorporates an external verifier to remove low-quality samples. Such verifiers may be human annotators or stronger teacher models. Filtering is widely used in iterative retraining and has shown empirical success in preventing model degradation and even improving performance (He et al., 2023; Tian et al., 2023; Guo et al., 2024; Zelikman et al., 2022; Zhang et al., 2024; Lampis et al., 2023; Haluptzok et al., 2023; Patwa et al., 2024). Motivated by this, we develop a principled understanding of when improvement is possible—namely, whether a generative model can leverage the verifier’s feedback, embedded in the selected synthetic subset, to achieve sustained gains.

Filtering and selecting synthetic data. While a rich line of empirical work demonstrates that these filtering strategies can improve model performance, theoretical understanding about iterative retraining with filtered synthetic data remains largely unexplored, with only a few recent exceptions. Amin et al. (2025) assume a strong, reliable quality function and focus on how an external labeler aids learning under this fixed filtering mechanism. Feng et al.

Method. In this section, we formalize our model of iterative retraining with verified synthetic data, coined verifier-based synthetic retraining for convenience. Following recent works in this space (Dohmatob et al., 2024a; Gerstgrasser et al., 2024; Garg et al., 2025; Zhu et al., 2025), we focus on the foundational linear regression setting where the objective is to estimate a high-dimensional coefficient vector θ⋆in the following linear model Modeling the verifier and data filtering rule. Suppose we have access to a verifier that possesses prior knowledge of θ⋆, modeled by a knowledge set. Specifically, the verifier’s knowledge is described by a spherical ball: Br(θc) := θ ∈Rp : ∥θ −θc∥≤r , with fixed center θc and radius r. We assume this knowledge set indeed contains the true parameter, i.e., θ⋆∈Br(θc), but the true parameter θ⋆is unknown. The verifier does not reveal θc or r directly (see modeling motivations below). Instead, it only provides binary feedback indicating whether a given (real or synthetic) data point (xi, yi) is consistent with the knowledge θ⋆∈Br(θc) or not. Specifically, the verifier outputs Yes if We refer to ∆= ∥θ⋆−θc∥as the bias of the verifier, whereas r captures the selectivity of the verifier – the smaller r is, less likely the verifier accepts a data point (xi, yi). The verifier only needs to provide Yes/No answers based on the above selection rule in equation 1, but does not needs to know the parameter θc, r of the knowledge set. The motivation of this modeling primarily comes from practice, as explained below.

Motivation of binary feedback from verifiers. We adopt the binary feedback from verifiers mainly for practical reasons. In practice, eliciting simple yes/no feedback is far less noisy and more cost-effective than asking verifiers to directly specify θc or r. Indeed, in real applications verifiers may not even know these quantities explicitly, which would correspond to model parameters if the verifier is a stronger teacher model or how the human reasons if the verifier is a human. This model choice is also aligned with the widely adopted comparison-based feedback in reinforcement learning from human feedback (RLHF) (Ouyang et al., 2022). Such binary feedback has become a standard approach in preference alignment for large language models, where LLM raters and human evaluators provide pairwise or accept/reject judgments that effectively guide learning at scale (Wettig et al., 2024). Although simple, our theory and empirical evaluations both show that this single bit of information for each sample can successfully be injected into the retraining process to improve models.

Synthetic Retraining with Verifier-based Filtering We begin with an initial set of real data (X0, Y 0), where X0 ∈Rn0×p and Y 0 ∈Rn0. The initial estimator ˆθ0 is obtained via Ordinary Least Squares (OLS) 4 Since learning proceeds through the conditional Y k | Xk, synthetic retraining requires specifying the covariate design Xk; labels Y k are then generated conditionally via the model under verifier constraints. In principle, one could construct Xk arbitrarily; however, for mathematical clarity, below we describe a targeted though arguably natural design. In particular, we choose to align the synthetic covariates with a fixed orthonormal set {v1, . . . , vp} and construct Xk in a block-structured form by repeating each v⊤ j as rows:

After verifier filtering, each orthogonal direction vj retains exactly nk samples with n0 ≤pn1 ≤pn2 ≤· · · .

Notably, the estimation of parameter ˆθk+1 using only synthetic data from the model with ˆθk, though with filtering, leads to a Markovian transition ˆθk 7→ˆθk+1. The above block design essentially helps “diagonalize” the transition operator ˆθk 7→ˆθk+1. The conceptual benefit of this covariance design choice is that we remove the rotational variability that arbitrary designs would introduce across iterations and decouple the dynamics along orthogonal directions. In practice, this design mirrors curating data along approximately orthogonal latent spaces or topics (e.g., topical axes like politics, sports, mathematics).

Discussion. The underlying mechanism is a verifier-induced bias-variance trade-off: filtering synthetic data reduces variance but may introduce bias.

Our key finding is that both behaviors can occur in long-term iterative retraining. The outcome depends critically on three factors: the growth rate of synthetic data, the verifier’s bias, and the verifier’s capability (i.e., its ability to reduce variance). Over time, iterative retraining injects increasingly more verifier knowledge into the estimator, while the contribution from the original data gradually decays. As a result, the verifier and the generative model family eventually dominate the limit behavior, driving the estimator ˆθk toward a fixed point, which corresponds to the verifier’s knowledge center θc.

This dynamic gives rise to three distinct phases of long-term behavior: (1) Unbiased verifier: If the verifier is unbiased (i.e., θc = θ⋆), iterative retraining yields continuous improvement and the estimator converges to the true parameter. (2) Mildly biased verifier: With small bias, iterative retraining can improve performance in the short term by reducing variance, but performance eventually plateaus or deteriorates as verifier bias accumulates. (3) Strongly biased verifier: With large bias, iterative retraining leads to degradation and may even cause collapse in the limit. Among these, the mildly biased case is the most practically relevant. It highlights a cautionary message: while synthetic retraining can initially boost accuracy, a perfectly unbiased verifier is unrealistic; consequently, this inherent bias will ultimately prevent sustained improvement.

Our study provides a theoretical and empirical characterization of verifier-guided synthetic retraining. We show that the process yields short-term gains by reducing variance through verifier filtering, but in the long run the estimator converges to the verifier’s knowledge center. This explains both the promise and the risk of such methods: a high-quality verifier can inject reliable external knowledge, while a biased verifier inevitably steers the model away from the truth. Viewed through the lens of information elicitation, our framework formalizes how external signals are incorporated recursively into training and why the outcome reflects the verifier’s information.

Limitations. Meanwhile, we also acknowledge the limitations of our results.Primarily, our analytical testbed relies on a well-specified parametric setting (linear regression) that assumes the existence of a global, ground-truth optimal parameter θ∗. This idealized assumption mathematically abstracts the complex, often competing attributes of a "good" generative model—such as sample diversity and generation quality—into a singular distance metric. In practice, while models like LLMs are parametric, they are approximating a true data-generating process that is unknown and highly likely non-parametric. For complex domains like natural language, a singular "true model" may not even exist; therefore, defining or evaluating an optimal model remains an open challenge. While our empirical extensions to VAEs and LLMs validate the theory qualitatively, formal generalization to richer models such as exponential families or simple neural network architectures are interesting future directions. Other future venues include developing sharper bounds for nonlinear models, exploring effectiveness of alternative synthetic design strategies beyond block orthogonalization, and studying verifier dynamics in large language models (LLMs) and vision models.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do models learn from self-generated outputs without cascading failures? How do training data quality and composition affect downstream model performance? How do interpretive frames override surface features in text comprehension? How do hallucinated citations emerge in AI scholarly output? Does training data format shape model reasoning more than domain content? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Why do training associations persist despite contradictory contextual information? How do curriculum design and feedback approaches affect model learning? How reliably can humans and AI detectors identify machine-generated text? How effectively can test-time voting aggregate diverse reasoning samples? Can confidence signals reliably detect flawed reasoning in language models? Why does AI verification capability persistently exceed generation capability?