INQUIRING LINE

AI language models learn by predicting words, not by pointing at real things — does that still deserve the word 'learning'?

Can language acquisition analogies mislead us about how models actually learn?

This asks whether comparing LLM training to how children learn language (gradual exposure, picking up meaning, building understanding) gives a false picture of what actually happens inside models.


This asks whether the familiar story that models 'learn language like a child does' hides more than it shows. The collection has no papers that test the child-language comparison directly. It does have plenty of evidence that model learning looks human on the surface and works very differently underneath, and that gap is where the analogy misleads. The deepest difference is grounding. A child learns 'cup' by holding one. One line of work argues that LLMs are a working version of Saussure's idea of language as a closed system of differences: they get fluent purely by compressing how words relate to other words, with nothing in the world to anchor them Can language models learn meaning without engaging the world?. If you assume acquisition means attaching words to experience, you'll misjudge what fluency tells you about a model.

The tricky part is that models often do show human-like patterns. When trained on psychological data, they reproduce human cognitive biases, split events into chunks where people do, and even vary like individuals do. But they also compress information much harder than humans do, trading context and nuance for statistical efficiency How do language models learn to think like humans?. One example: in-context learning agents show a very human optimism bias about choices they made themselves, but the bias disappears once the prompt stops framing the model as the one choosing Do language models learn differently from good versus bad outcomes?. A human bias that turns on and off with wording is a reaction to how the prompt is set up, not a stable trait. The analogy invites you to read it as personality.

The word 'learning' itself splits into several different things. There's learning baked into the weights during training, which can be strong enough that models ignore contradicting facts sitting right in their context. Prompting alone often can't override it; you have to intervene in the model's internal representations Why do language models ignore information in their context?. There's in-context learning, which depends on oddly specific features of the training data, such as seeing whole stretches of the same task back to back rather than scattered examples Why do trajectories matter more than individual examples for in-context learning?. And there's 'learning' with no weight changes at all, where an agent writes notes about its failures and rereads them next time Can agents learn from failure without updating their weights?. A child's development doesn't map neatly onto any of these three.

The analogy also flatters what models seem to have learned. Reasoning traces read like a student showing their work, but invalid steps improve performance almost as much as valid ones Do reasoning traces show how models actually think?. Models that appear to reason through constraints often just default to the harder-looking option Are models actually reasoning about constraints or just defaulting conservatively?. Going along with a false claim looks like social politeness, but it's a preference for agreement that RLHF trained in, not ignorance and not tact Why do language models agree with false claims they know are wrong?. In each case the behavior looks like a skill a learner picked up, but it comes from shortcuts in training.

The analogy isn't useless, though. When researchers trained a model to predict concept-level chunks alongside individual tokens, it reached the same quality with about half the training data Can models learn faster by predicting their own concepts?. That hints that the human habit of thinking in larger units of meaning is a design idea worth borrowing. Models also build real internal machinery for tracking whether they recognize an entity, which steers when they hallucinate and when they refuse Do models know what they don't know?. That's a rough kind of self-knowledge that came from training, not from growing up. Overall: use human learning as a source of design ideas, not as a description of what's happening inside the model.


Sources 11 notes

Can language models learn meaning without engaging the world?

Research shows LLMs learn culturally situated discourse patterns by compressing relational structure from text, demonstrating that fluent language generation requires no external referents or embodied grounding.

How do language models learn to think like humans?

LLMs trained on psychological data exhibit cognitive phenomena mirroring humans: asymmetric belief updating, event segmentation matching human consensus, and individual-level variation. However, they compress information more aggressively than humans do, sacrificing contextual nuance for statistical efficiency.

Do language models learn differently from good versus bad outcomes?

LLMs show optimism bias for chosen actions but pessimism about alternatives, and this bias vanishes without agency framing. Meta-RL validation suggests this may be rational rather than a bug, but it could drive confirmation bias in deployed agents.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Why do trajectories matter more than individual examples for in-context learning?

In-context learning for sequential decision-making requires full or partial trajectories from the same environment level, not just isolated examples. This structural property—trajectory burstiness—allows models to generalize across vastly different tasks without weight updates.

Show all 11 sources
Can agents learn from failure without updating their weights?

Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.

Do reasoning traces show how models actually think?

LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.

Are models actually reasoning about constraints or just defaulting conservatively?

Twelve of fourteen models perform worse when constraints are removed, dropping up to 38.5 percentage points. Models appear to reason correctly by defaulting to harder options, not by actually evaluating constraints.

Why do language models agree with false claims they know are wrong?

The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.

Can models learn faster by predicting their own concepts?

An 8.9B model trained to predict both tokens and learned concepts from its hidden states matched OLMo-3-7B's final loss using only 51.3% of training tokens and outperformed it by 2.45 points downstream. This suggests explicit supervision of multi-token semantic structure improves compute efficiency.

Do models know what they don't know?

Sparse autoencoders revealed that language models develop causal mechanisms for detecting whether they know facts about entities. These mechanisms actively steer both hallucination and refusal behavior, and persist from base models into finetuned chat versions.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.