When an AI seems to have a new idea, is it truly its own, or inherited from its training data?
How do current models inherit concepts from training data rather than initiate them?
This explores how much of what a language model 'knows' or 'thinks with' comes from patterns already in its training data, and whether there are any signs of models forming their own concepts instead of just recombining inherited ones.
This explores how much a model's ideas are inherited from its training data versus formed by the model itself. The short version from the corpus: inheritance runs deeper than most people expect. Even methods that look like they teach models new things often just pick out, amplify or reorganize what pretraining already put there. A handful of newer approaches try to give models a way to form their own concepts, but even those mostly work by having the model predict structure it has already absorbed.
The clearest evidence of inheritance is how strongly training-data associations resist change. When what a model learned in training conflicts with what you tell it in the prompt, training often wins. Rewording the prompt isn't enough, and researchers had to reach into the model's internal representations to make it use the new information Why do language models ignore information in their context?. Reinforcement learning, often described as how models 'learn to reason,' shows the same pattern from another angle. In controlled experiments, RL quickly latches onto one answer format that already existed in the pretraining data and suppresses the others, usually within the first pass over the data. Which format wins depends on model size, not on which format works best Does RL training collapse format diversity in pretrained models?. So RL acts less like a teacher and more like a filter on what the model already contains.
The inheritance is so strong that some of the most effective methods are built around protecting it. Proxy-tuning steers a model's outputs at generation time without touching its weights, and it beats direct fine-tuning on knowledge tasks, because fine-tuning damages where knowledge is stored in the lower layers Can decoding-time tuning preserve knowledge better than weight fine-tuning?. Editing hidden representations instead of weights gets similar or better results with a small fraction of the parameters Can editing hidden representations beat weight updates for finetuning?. Models that stay statistically closer to their base model also keep more ability to learn the next task Does staying close to the base model preserve learning ability?. Read together, these suggest the pretrained base is the real store of concepts, and post-training mostly reaches into it.
On the 'initiate' side, the corpus has suggestive work but nothing conclusive. One model was trained to predict higher-level concepts drawn from its own hidden states alongside the next word. It matched a comparable model's final loss using about half the training data Can models learn faster by predicting their own concepts?. A formal analysis explains why: predicting your own internal summaries recovers layered structure exponentially faster than predicting raw words Why is predicting latents more sample-efficient than tokens?. Other work has models produce their own intermediate material, such as writing practice problems for themselves Can language models improve themselves without any external training data? or deliberately making mistakes and then stating the general principle behind them Does learning from mistakes improve in-context learning?. Look closely, though, and each one still builds on what the model already has: its own hidden states, its own judgment of what counts as a good problem, its own existing knowledge.
The twist you might not expect: how knowledge is organized may matter more than how much there is. Arranging training text into an auto-generated topic hierarchy, so the model learns where each fact sits within a conceptual map, got half of full-training performance from 0.3% of the data Can organizing knowledge structures beat raw training data volume?. That points to a middle ground between inheriting and initiating: models may not invent concepts, but they can be much better or worse at arranging the ones they inherit. The corpus doesn't yet have direct work on models forming truly new concepts, so treat the 'initiate' half of this question as open.
Sources 10 notes
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Proxy-tuning closes 88-91% of the alignment gap while surpassing direct fine-tuning on knowledge tasks by leaving base model weights untouched. Direct fine-tuning corrupts knowledge storage in lower layers, whereas proxy-tuning applies distributional shifts that primarily affect reasoning and style.
ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.
FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.
Show all 10 sources
An 8.9B model trained to predict both tokens and learned concepts from its hidden states matched OLMo-3-7B's final loss using only 51.3% of training tokens and outperformed it by 2.45 points downstream. This suggests explicit supervision of multi-token semantic structure improves compute efficiency.
A formal sample-complexity analysis proves latent-level self-supervision (data2vec/JEPA style) recovers compositional structure with samples constant in hierarchy depth, while token-level learning requires exponential samples—because same-level latents are far more correlated than raw tokens.
SQLM uses a proposer-solver framework where the proposer generates calibrated problems and the solver learns via majority-vote verification. Both agents improve through RL alone, creating an automatic curriculum that scales without human labels or ground-truth answers.
LEAP demonstrates that models achieve better performance on reasoning and math tasks by intentionally erring on few-shot examples, reflecting on mistakes, and deriving explicit task-specific principles—without additional labeled data or fine-tuning.
StructTuning achieves 50% of full-corpus performance using only 0.3% of training data by organizing chunks into auto-generated domain taxonomies. The model learns knowledge position within conceptual structures rather than raw text patterns, matching how students learn from textbooks.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Learn from your own latents and not from tokens: A sample-complexity theory
- NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- How new data permeates LLM knowledge and how to dilute it
- Learning To Retrieve Prompts for In-Context Learning
- Sharpening Tax in Post-Training
- Learning, Fast and Slow: Towards LLMs That Adapt Continually
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs