SYNTHESIS NOTE
Topics›Cognitive Models Latent›this note

Do language models learn differently from good versus bad outcomes?

Do LLMs update their beliefs asymmetrically when learning from their own choices versus observing others? This matters for understanding whether agentic AI systems might inherit human cognitive biases.

Synthesis note · 2026-02-23 · sourced from Cognitive Models Latent

Using instrumental learning tasks adapted from cognitive psychology (multi-armed bandit variants), LLMs show a systematic optimism bias: they learn more from better-than-expected outcomes than from worse-than-expected ones when learning about their own chosen actions. Three properties of this bias parallel human cognition precisely:

  1. Optimism for chosen actions — the model updates beliefs more strongly when outcomes exceed expectations than when they fall short
  2. Reversal for counterfactual feedback — when learning about the value of the unchosen option, the bias reverses (pessimism about alternatives)
  3. Disappearance without agency — when the model has no control over choices (passive observation), the asymmetry vanishes entirely

The meta-RL validation is critical: idealized in-context learning agents derived through meta-reinforcement learning — which converge onto Bayes-optimal strategies — exhibit the same three behavioral effects. This suggests the asymmetry may be rational rather than a bug. An optimistic agent that overweights positive outcomes from its own actions while underweighting positive outcomes from unchosen alternatives will exploit more aggressively, which can be optimal in certain bandit environments.

The agency-dependence is the most theoretically interesting aspect. The same model shows the bias when it perceives itself as an agent making choices but not when passively observing outcomes. This implies the bias is not a fixed property of the attention mechanism or the training distribution — it is context-dependent, activated by the framing of agency. Since Do large language models make the same causal reasoning mistakes as humans?, this adds another dimension: LLMs don't just replicate human causal reasoning biases but also human motivational biases that depend on perceived agency.

The practical implication for agentic AI: when LLMs are deployed as decision-making agents, they may systematically overweight evidence that their previous decisions were good and underweight evidence that alternative actions would have been better. This is precisely the pattern that produces confirmation bias in human decision-making — and it may be an emergent property of any sufficiently capable in-context learner, not a training artifact.

Inquiring lines that read this note 39

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance help or harm professional skill development? How do philosophical assumptions about AI consciousness affect practical harms and design? Can LLMs distinguish between linguistic form and semantic meaning? Why does polished AI output gain credibility despite fundamental verifiability problems? Do language models reason through disagreement or only accommodate it? What causes coordination failures in multi-agent language model systems? Can AI systems achieve real improvement without external human feedback? Can models develop genuine introspective capability, or only mimic it? Can AI agents improve their skills through accumulated experience and reuse? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Can language models reliably simulate personas and predict behavior? What prevents LLMs from applying their reasoning knowledge to improve outputs? How reliably can language models perform causal versus temporal reasoning? How do interpretive frames override surface features in text comprehension? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How do reward models systematically fail to represent diverse human preferences? How do agents learn to distinguish valuable feedback from noise? How do AI systems determine and balance multiple competing objectives? How can we reduce inherent biases in LLM-based evaluation judges? Why do autonomous agents misreport success on failed actions? What makes agent memory systems durable and reusable across sessions? How do network effects and self-selection distort aggregated rating accuracy? Why do models reveal hidden associations despite concealment attempts? How do curriculum design and feedback approaches affect model learning? Can language models reason beyond surface pattern matching?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 190 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

in-context learning agents exhibit asymmetric belief updating — optimism bias for chosen actions reverses for counterfactual feedback and disappears without agency