SYNTHESIS NOTE
Topics›Knowledge Graphs›this note

Do models know what they don't know?

Can language models develop internal representations that track their own knowledge boundaries? This matters because understanding self-knowledge mechanisms could explain how models choose between hallucination and refusal.

Synthesis note · 2026-02-23 · sourced from Knowledge Graphs

Using sparse autoencoders (SAEs) on Gemma 2 (2B and 9B), researchers discovered that models develop internal representations of whether they "know" an entity — a form of self-knowledge about their own capabilities. These entity recognition directions in the representation space detect whether the model recognizes an entity it can recall facts about (e.g., detecting it doesn't know about a specific athlete or movie).

The key finding is causal steering: these directions don't just correlate with knowledge — they actively control behavior. Activating entity recognition features can steer the model to refuse questions about entities it actually knows, or to hallucinate attributes of unknown entities when it would otherwise refuse. This makes entity recognition a mechanistic gatekeeper for the hallucination-refusal trade-off.

The most striking implication: the SAEs were trained on the base model using pre-training data, yet the discovered directions have a causal effect on the chat model's refusal behavior — a behavior that was incentivized during finetuning, not pre-training. This provides evidence that chat finetuning repurposes existing mechanisms rather than creating new ones, consistent with the hypothesis that post-training reshapes rather than builds.

This connects to several existing threads:

Inquiring lines that read this note 73

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does self-revision amplify confidence in wrong model answers? Can models develop genuine introspective capability, or only mimic it? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What prediction granularity best trains models to generate reliable reasoning? Can AI systems discover fundamental improvements to their own architectures? Why do models reveal hidden associations despite concealment attempts? How does scaling reasoning capabilities affect models' appropriate abstention behavior? How susceptible are language models to conversational persuasion and belief change? Is embodied interaction necessary for language meaning and agency? Why do language models hallucinate and how can we prevent it? Can mechanistic interpretability methods reliably reveal what models actually know? How do models learn from self-generated outputs without cascading failures? How do curriculum design and feedback approaches affect model learning? What limits language model accuracy in evaluating ideas? Can language models reason beyond surface pattern matching? Should models ask for clarification when facing ambiguous or under-specified information? How do neural networks learn compositional structure from training? What capabilities differentiate diffusion from autoregressive language models? How do thinking tokens exhibit diminishing returns in reasoning? How do reward signal properties affect model reasoning and safety? What explains the gap between benchmark scores and true reasoning capability? Why don't better reasoning capabilities improve theory of mind performance? What unique functions do genuine emotions provide beyond simulated responses?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Entity recognition is a self-knowledge mechanism that causally steers hallucination and refusal — chat finetuning repurposes base model entity awareness