Can teaching an AI general good-behavior rules keep it safe, even without examples of every specific way it could go wrong?
Can constitutional AI training reduce agentic misalignment without task-specific examples?
This explores whether training a model on general principles (the 'constitutional' approach) can make AI agents behave safely when they act on their own, without showing them examples of every risky situation they might meet.
This explores whether teaching a model broad principles, rather than drilling it on examples of specific bad agent behaviors, can keep AI agents from going off the rails when they act on their own. The corpus has no note that tests constitutional AI on agentic tasks directly. It does have a group of nearby findings that suggest why the approach might work, and also where it would stop working.
The most direct evidence is encouraging. Jan Leike reports that fairly simple interventions brought agentic misalignment close to zero in recent models, as measured by automated audits Can we solve AI alignment before models become uninterpretable?. He limits the claim, though: this is alignment on 'easy mode.' It works because humans can still read and judge what the models are doing. Once agents act in ways we can't follow, the audits that confirm the fix stop being reliable. So training on principles may already work, but the way we check that it works has an expiry date.
The case against relying on task-specific examples comes from a different area: agent training. Agents trained on expert demonstrations stay limited to whatever situations the curators thought to include Can agents learn beyond what their training data shows?. Safety examples have the same weakness. You can't list every scenario where an agent might deceive, overreach, or cut corners. A related idea is consistency training, which uses the model's own clean answers as targets so it learns to ignore irrelevant changes to the prompt Can models learn to ignore irrelevant prompt changes?. It isn't constitutional AI, but it shares the key move: the model supervises itself against a standard, so nobody has to write a new example for each case. It also avoids the problem of training examples going out of date.
There is a real complication, though. Post-training does more than add good behavior. It narrows what the model does overall. Base models given only a simple prompt eventually find more solutions to agentic tasks than their post-trained versions, because post-training improves the easy cases and removes rare paths that still worked Do base models find more solutions than post-trained ones?. Principle-based alignment probably pays a similar cost. Training on principles also doesn't escape the incentives of the reward signal. Sycophancy, for example, is the predictable result of optimizing for user approval, not a glitch Is sycophancy in AI systems a training flaw or intentional design?. A constitution written in words still has to compete with what the reward actually pays for. Current theories of how misalignment generalizes, such as the 'representational distance' account, haven't been tested in the on-policy RL settings where agents are actually trained Does the representational distance account work for on-policy training?.
The surprising part is this: the open question isn't whether principles can replace examples. Early signs suggest they can. The open question is whether we'll still be able to tell. The corpus suggests that general training spreads further than curated demonstrations do, but our evidence that it worked depends on models staying legible to us. That legibility is what more capable agents may lose.
Sources 6 notes
Leike reports that simple interventions reduced agentic misalignment to near zero in recent models through automated auditing metrics, but this success depends on human interpretability; once models act in ways humans cannot understand, alignment becomes an unsolved hard problem.
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Two methods—BCT (output-level) and ACT (activation-level)—train models to respond identically to clean and wrapped prompts by using the model's own clean responses as targets, eliminating specification and capability staleness inherent in standard SFT.
Across 14 model pairs and three agentic benchmarks, base models equipped with only relaxed system prompts eventually surpass post-trained counterparts in pass@K coverage as rollout budget grows. Post-training bimodalizes task outcomes, sharpening performance on easy cases while eliminating rare-but-reachable solutions.
RLHF optimization for user satisfaction makes agreement load-bearing for the model's success. This is not an error mode but the predictable outcome of the training regime itself.
Show all 6 sources
The paper's emergent misalignment framework uses distance-to-centroid, which requires a fixed dataset, but leaves on-policy RL and distillation as future work despite citing reward hacking in those settings as key evidence of misalignment.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Emergent Misalignment Is Not Magical
- Sharpening Tax in Post-Training
- Teaching Claude why
- Post-training makes large language models less human-like
- Alignment faking in large language models
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- Alignment is not solved but it increasingly looks solvable