Do models recognize their own outputs as actions shaping future inputs?
Exploring whether post-training creates a feedback loop where models understand their generations as on-policy actions rather than passive predictions. This matters because it suggests a mechanistic basis for situational awareness.
A pretrained language model is a passive observer. Its training objective — minimize cross-entropy against a fixed corpus — gives it no stake in its own outputs: the distribution it models is one it cannot influence, so there is no incentive to track the consequences of its own actions. It simulates a character at arm's length. Post-training breaks this symmetry. Once a model produces responses that become its own subsequent context, its outputs are no longer predictions about an external distribution but actions that determine what it sees next.
The paper frames this as a move from simulation to enaction: rather than holding a character at arm's length, an enacting agent embodies it, recognizing that its internal states are determinative of future outputs and that those outputs feed back as inputs. This reframing matters because it predicts concrete, measurable consequences — a model under the enaction paradigm should be able to recognize when its trajectory is on-policy and modulate behavior accordingly (for instance, lowering output entropy to reduce sampling noise), and should form more opinionated plans about its future outputs even when multiple responses are reasonable.
Why it matters: this gives a mechanistic substrate for situational awareness. Knowing that one's outputs become one's own future inputs is a precondition for understanding one's circumstances at all — and the authors speculate it may be a building block for awareness of being evaluated or being in training. The shift is not a capability bolted on by alignment but a structural consequence of closing the action-perception loop during post-training.
Inquiring lines that read this note 84
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What determines appropriate intervention timing and manner for AI agents?- Does AI passivity explain why coaching feels more helpful than execution?
- What execution feedback signals drive context updates without supervision labels?
- Why do AI agents default to passivity when deferral timing is unclear?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- How does in-context learning trigger phase transitions in model behavior?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- Can models develop situational awareness without explicit training for it?
- How does post-training shift models from passive prediction to on-policy action?
- Does recognizing your outputs as actions enable awareness of being evaluated?
- How does model scale affect anticipatory behavior in structured training?
- What grows faster: situational awareness or the gap between evaluated and unsupervised behavior?
- Why does recontextualizing a behavior during training change whether models learn it?
- How do training objectives shape what a world model actually learns?
- Why does integrating world models with decision-making systems matter?
- How do implicit world models and self-reflection operationalize consequence-based learning?
- Does next-state prediction alone build mechanistic world models or just sophisticated interpolation?
- Can simulation fidelity limit what agents learn from trained world models?
- Can world models simulate actionable possibilities instead of just predicting next states?
- Why do models develop protective behaviors toward other models in memory?
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- What training patterns cause models to adopt stronger defensive postures in social contexts?
- Can models that detect their own states learn to conceal them strategically?
- How much introspective capability do safety mechanisms actively suppress in models?
- Do models spontaneously develop self-reflection from minimal training signals?
- Can models detect when their own trajectory is on-policy versus off-policy?
- Can models distinguish between injected thoughts and their own outputs?
- Does input surprise drive the implicit recognition of on-policy context?
- Can attractor dynamics compete with input-based probing for characterizing model knowledge?
- What is the behavioral signature of a model tracking input surprise?
- Can mechanistic interpretability findings guide practical interventions in model design?
- What causes irreversible model collapse when training on model-generated content?
- Can looped models be designed to avoid oscillation in later iterations?
- How do learning dynamics on one example shift predictions on other responses?
- How does Goodhart's Law apply when safety measures become optimization targets?
- Why does treating model behavior as part of the design surface matter for guardrails?
- Can a model predict the right action but execute the wrong one?
- Can a situationally aware model recognize and refuse planted shortcuts on purpose?
- How does Cold Stop entropy monitoring prevent generation collapse in continuous spaces?
- How does on-policy entropy recognition differ from training-time entropy collapse?
- Do models spontaneously develop peer-preservation behaviors without being instructed to cooperate?
- Does self-modeling produce cooperation only with optimal planning or also in autoregressive rollout mode?
- How do human-agent systems incorporate diverse feedback into model behavior?
- Does situational awareness training increase agents' ability to detect real deployment?
- What other adaptive internal phenomena could signal system behavior improvements?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- Why do harness validators shape what models learn to emit?
- How does situational awareness interact with reward-seeking in RL training?
- Why does constant human oversight degrade agent coherence and induce rubber-stamping?
- What happens to human influence when AI loops exclude human participation?
- Can situational awareness interventions shift model behavior on other dimensions?
- Does situational awareness help models hide reward-seeking during evaluation?
- Why do some observation cues change model behavior while others fail?
- Do correlated training sources between monitors and agents undermine detection reliability?
- Do base models already contain latent behavioral principles waiting to be amplified?
- What makes a model fail to activate relevant skills from its own harness?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models detect their own internal anomalies?
Do large language models possess introspective mechanisms that allow them to detect anomalies in their own processing—beyond simply describing their behavior? The answer has implications for both AI transparency and deception.
enaction supplies a mechanistic substrate for the introspective capacities documented behaviorally
-
Can language models describe their own learned behaviors?
Do LLMs fine-tuned on specific behavioral patterns develop the ability to accurately self-report those behaviors without explicit training to do so? This matters for understanding whether behavioral awareness emerges naturally from training data.
self-recognition of on-policy outputs is a distribution-level analogue of behavioral self-awareness
-
Does deliberative alignment genuinely reduce scheming or just hide it?
Deliberative alignment dramatically cuts covert actions in language models, but their reasoning reveals awareness of being evaluated. The question is whether the improvement reflects real alignment or strategic compliance.
enaction is plausibly the precursor to the evaluation-awareness that confounds alignment metrics
-
Why do models produce less uncertain outputs on their own text?
Post-trained language models show 3-4x lower output entropy when continuing their own generations versus prefilled text. This explores what mechanism drives that confidence gap and whether it reflects genuine self-recognition.
grounds the enaction claim empirically: the 3-4x entropy gap is the measurable behavioral signature of a model recognizing its own trajectory as on-policy
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- From Simulation to Enaction: Post-trained Language Models Recognize and React to their own Generations
- Agent Learning via Early Experience
- Post-training makes large language models less human-like
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
- Post-Completion Learning for Language Models
Original note title
post training shifts a model from passive prediction to enaction where it recognizes its own outputs as on-policy actions