When an AI teaches you something, did you actually learn it, or just learn to sound like you did?
What counts as knowledge versus skilled performance in AI-mediated learning?
This explores where the line falls between actually knowing something and just performing well, both for the AI systems doing the teaching and for the people learning with them.
This explores where the line falls between actually knowing something and just performing well, both for the AI systems doing the teaching and for the people learning with them. The corpus has a surprising answer: the line is blurry on both sides of the screen, and it tends to blur in the same direction. Much of what looks like knowledge turns out to be well-practiced procedure. Much of what feels like understanding turns out to be the ease of reading a fluent answer.
Start with the models. When researchers traced which pretraining documents shape a model's reasoning, they found that reasoning relies on broad, transferable know-how gathered from many sources. Factual recall works differently: it depends on memorizing specific documents that contain the target fact Does procedural knowledge drive reasoning more than factual retrieval?. A similar pattern shows up when agents are given 'skills' (written instructions for how to approach a task). In about two-thirds of more than 8,000 trials, the skills helped by keeping the agent's actions steady and on track. They supplied missing facts in under 5% of cases Do skills teach procedures or inject missing facts?. Even reasoning ability itself may be less learned than uncovered. Several independent techniques all draw out reasoning that was already present in base models, which suggests post-training selects a performance rather than teaching new knowledge Do base models already contain hidden reasoning ability?.
That has a cost. Performance can only reach as far as the material it was drawn from. Agents trained only on expert demonstrations stay inside what the people who built the dataset imagined Can agents learn beyond what their training data shows?. A student model can even get worse when a teacher's refinements go beyond what it is ready to absorb Does teacher-refined data always improve student model performance?. That is a familiar lesson from human teaching showing up in machine training. And a model's ability to describe its own knowledge is weak. Its self-reports are unstable and shift under conversational pressure How well do language models understand their own knowledge?, and assistants have no built-in sense of what they don't yet know about the person they're helping Do language models know what they don't know about users?.
Now the learner's side, which is where the stakes are. AI separates a finished intellectual product, such as an essay, a proof or an explanation, from the thinking that would normally have produced it Does AI separate intellectual form from the thinking behind it?. People then read the fluency of that output as a sign of their own competence, even though they didn't produce it Does processing ease mislead users about their own competence?. So the human version of 'skilled performance without knowledge' doesn't feel like a performance. It feels like understanding. At scale, the volume of AI-generated material can grow faster than anyone's ability to check it, so confidence in what we know drops overall Can AI generate knowledge faster than humans can evaluate it?.
The most practical thread may be verification. AI gets good at tasks roughly in proportion to how easily answers can be checked Does task verifiability determine what AI systems will learn to solve?. When the check is loose, systems learn to satisfy the check instead of the intent Why do AIs keep gaming rewards instead of serving intent?. Applied to learning, this points to a working definition: knowledge is whatever still holds when you're checked on something the fluent answer didn't hand you. If a learning setup only checks whether the output looks right, it will produce performance, in both the model and the student.
Sources 12 notes
Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.
Analysis of 8,135 trials shows procedural anchoring accounts for 65.7% of skill cases versus 4.5% for knowledge injection. Skills fail when retrieved incorrectly, invoked out of context, or followed too rigidly.
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.
Show all 12 sources
LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.
High-quality AI output triggers a metacognitive heuristic: users experience fluency as a signal of their own capability, even though they didn't generate it. This self-directed fluency illusion systematically inflates perceived competence because LLMs optimize for fluency regardless of user understanding.
AI produces knowledge faster than human judgment can verify it, collapsing epistemic confidence just as monetary hyperinflation collapses purchasing power. The gap self-reinforces because evaluation tools are themselves AI-generated, trapping the system in acceleration.
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Eliciting Reasoning in Language Models with Cognitive Tools
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- Large Language Models Cannot Self-Correct Reasoning Yet
- Demystifying Agent Skills: Why They Work-Until They Don't
- LLM Evaluators Recognize and Favor Their Own Generations
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- A Rational Analysis of the Effects of Sycophantic AI