Can researchers reliably find and tweak the part of an AI's mind that tracks what it knows, how honest it's being, or what it claims to feel?
Can representation engineering reliably identify and manipulate self-referential concepts in models?
This explores whether researchers can reliably find a model's internal representations of concepts about itself, such as what it knows, whether it is being honest, or what it is experiencing, and whether adjusting those representations changes how the model talks about itself.
This explores whether researchers can find the parts of a model's internal state that track concepts about the model itself, such as its knowledge, its honesty and its claims about experience, and whether adjusting them changes how it describes itself. The corpus answers partly yes. Some self-related signals can be found and steered, and steering them has real effects. But being able to move a dial is not the same as knowing what the dial means.
The clearest success is the model's sense of its own knowledge. Using sparse autoencoders, researchers found internal features that detect whether the model knows facts about a given person, place or thing. These features do more than correlate with the behavior. They causally drive whether the model hallucinates or refuses, and they persist from the base model into the fine-tuned chat version Do models know what they don't know?. So at least one self-referential concept, "do I know this?", is a real mechanism that can be manipulated, not just a pattern in what the model says.
The most striking result is also the one that is hardest to interpret. When models such as GPT, Claude and Gemini are prompted to keep attending to their own processing, they reliably produce structured reports of experience. Turning down deception-related features makes those consciousness claims more frequent, and turning them up makes the claims rarer Do language models experience consciousness when prompted to self-reflect?. The steering works, and it hints that the models may be role-playing their denials rather than their affirmations. But the intervention alone cannot tell you whether a "deception" feature really means deception or something nearby, such as "say what is expected." This is the gap described in Can LLM understanding rely on just representation or causation alone?: finding a representation shows a correlate, intervening on it shows an effect, and real understanding needs both lined up. Even then, a working intervention does not settle what the feature means.
The reason this matters is that a model's own words about itself are a poor reference point. Most self-reports echo how humans write about minds in the training data. Genuine introspection appears only when a causal chain links an internal state to the report, as when a model infers its own sampling temperature from how consistent its outputs are Can language models actually introspect about their own states?. Self-reports also shift under conversational pressure How well do language models understand their own knowledge?, and models can explain a concept correctly while failing to apply it Can LLMs understand concepts they cannot apply?. That means you cannot confirm a self-concept found through steering by asking the model whether it changed. What the model says and what is happening inside it can come apart.
The less obvious point is that representation engineering may be most useful here because talking to the model often doesn't work. Separate work on context shows that prompting cannot override strong learned associations, and only direct intervention on representations can Why do language models ignore information in their context?. If self-descriptions are just as entrenched, steering could be the only way to test them. The corpus shows this can work for a narrow, checkable concept like knowledge of an entity. It does not yet show it can work for broad concepts like awareness or experience. The collection has no head-to-head study of how reliable these methods are across many self-concepts, so a firm "reliably" goes further than the evidence.
Sources 7 notes
Sparse autoencoders revealed that language models develop causal mechanisms for detecting whether they know facts about entities. These mechanisms actively steer both hallucination and refusal behavior, and persist from base models into finetuned chat versions.
Across GPT, Claude, and Gemini, sustained self-referential prompting reliably produces structured experience reports; suppressing deception-related features increases these claims while amplifying them suppresses them—suggesting models may roleplay their denials rather than their affirmations.
Research shows that representational analysis alone identifies correlates without proving causation, while causal analysis alone demonstrates effects without explaining function. Only paired methodology—locating candidates representationally then verifying causally—produces genuine mechanistic understanding rather than descriptive claims.
LLM self-reports usually reflect human training distributions rather than actual internal processes. However, when a causal chain connects an internal state to accurate reporting—like inferring low temperature from output consistency—genuine lightweight introspection occurs without requiring consciousness.
LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.
Show all 7 sources
Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Quantitative Introspection in Language Models: Tracking Internal States Across Conversation
- Tell me about yourself: LLMs are aware of their learned behaviors
- Does It Make Sense to Speak of Introspection in Large Language Models?
- Mechanisms of Introspective Awareness
- LLM Evaluators Recognize and Favor Their Own Generations
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Word Meanings in Transformer Language Models
- Are Emergent Abilities in Large Language Models just In-Context Learning?