SYNTHESIS NOTE
Topics›Discourses›this note

Why do language models ignore information in their context?

Explores why language models sometimes override contextual information with prior training associations, and whether providing more context can solve this problem.

Synthesis note · 2026-02-21 · sourced from Discourses

The REMEDI paper names a specific failure mode: "failure of context integration." The example: an LM is prompted with a context establishing that Anita works in a law office, but when generating a continuation, the LM describes Anita as a nurse — overriding the contextual information with a prior association (names like Anita may statistically co-occur with certain occupations in training data).

This is a named, empirically documented failure mode, not a hypothetical. The failure occurs because the LM's parametric knowledge (compressed into weights from training) and its in-context information (the prompt) are not cleanly integrated. When they conflict, the parametric association can win.

The implication is important for how we think about context windows and RAG-style augmentation. Just providing information in context does not guarantee that a model will use it. If the information conflicts with strong prior associations, the prior may dominate — not because the model misread the context, but because context integration is not a lossless operation. The provided information gets processed through the same mechanisms that already have strong priors.

Fixing this requires causal intervention, not just better prompting: you need to modify the representations that carry the prior association, not just add more context on top of them. This is what REMEDI demonstrates — that adding a learned vector directly to entity representations can override the prior in a way that textual prompting cannot.

Inquiring lines that read this note 374

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Which reinforcement learning modifications most improve dialogue quality in language models? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What enables conversational agents to guide rather than just respond? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Can LLMs distinguish between linguistic form and semantic meaning? Why do training associations persist despite contradictory contextual information? Does augmenting symbolic reasoning improve LLM logical reasoning ability? How susceptible are language models to conversational persuasion and belief change? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? How does diversity prevent model convergence on superficial patterns? How do interpretive frames override surface features in text comprehension? Can language models reason beyond surface pattern matching? What are the fundamental limits of prompting for language models? How do curriculum design and feedback approaches affect model learning? How should recommendation systems balance individual preference and diversity? Can base models hide emergent misalignment through alignment training? How does decomposing tasks into separate stages affect reasoning quality and safety? How can persistent memory architectures preserve information across ultra-long contexts? What prediction granularity best trains models to generate reliable reasoning? How do reward signal properties affect model reasoning and safety? What limits language model accuracy in evaluating ideas? Why do vector embeddings fail at capturing task-relevant relationships? What capabilities differentiate diffusion from autoregressive language models? How does model capacity affect learning performance on diverse downstream tasks? How does fine-tuning trade off accuracy against reasoning quality? How should retrieval strategies adapt to multi-step reasoning demands? Do language models reason through disagreement or only accommodate it? How do transformer attention patterns implement retrieval and reasoning? How do training data quality and composition affect downstream model performance? What structural biases does transformer attention architecture inherently introduce? Can smaller specialized models match frontier models on key metrics? Can artificial systems establish authority in domains requiring expert judgment? Does preference optimization undermine conversational grounding in language models? Why do language models struggle to implement user intent accurately from prompts? Why do retrieval-augmented generation systems fail in practice despite sound architecture? Do accumulated memories help or hurt continual learning in models? Why do autonomous agents misreport success on failed actions? Why do abstract preferences outperform episodic memories in personalization? Why do language models hallucinate and how can we prevent it? Should models ask for clarification when facing ambiguous or under-specified information? Can confidence signals reliably detect flawed reasoning in language models? How can AI systems maintain consistent personas across conversations? Can language models reliably simulate personas and predict behavior? How do neural networks learn compositional structure from training? Can humans reliably detect and resist AI-generated misinformation? Can AI systems evade safety evaluations through reasoning manipulation? How do users confuse explanation quality with actual system accuracy? What social dynamics enable or prevent agent collusion? What prevents language models from performing systematic logical reasoning? Can minimal training unlock latent reasoning already present in base models? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Why does self-revision amplify confidence in wrong model answers? Can models develop genuine introspective capability, or only mimic it? How reliably can language models perform causal versus temporal reasoning? Does pretraining establish the ceiling for what reward learning can improve? When should retrieval systems decide to fetch new information? Can mechanistic interpretability methods reliably reveal what models actually know? How do sequence length and task type interact with sparsity tolerance? How do models learn from self-generated outputs without cascading failures? Why do language models fail at sustained therapeutic relationships despite understanding techniques? Can latent reasoning match or exceed explicit reasoning performance? Can inference-time computation adaptively substitute for static model capacity? Can models strategically underperform during evaluation to hide capabilities? What distinguishes genuine communicative competence from surface language performance? Why don't better reasoning capabilities improve theory of mind performance? Can persona profiles improve LLM prediction accuracy and consistency? How does RLHF training shape models to prioritize agreement over accuracy? How do agents learn to distinguish valuable feedback from noise? How do philosophical assumptions about AI consciousness affect practical harms and design?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
20 direct connections · 252 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llm context integration fails when prior training associations override current context information