Toward understanding and preventing misalignment generalization
Large language models like ChatGPT don’t just learn facts—they pick up on patterns of behavior. That means they can start to act like different “personas,” or types of people, based on the content they’ve been trained on. Some of those personas are helpful and honest. Others might be careless or misleading.
Existing research showed that if you train a model on wrong answers, even in just one narrow area, like writing insecure computer code, it can inadvertently cause the model to act “misaligned” in many other areas. This is called “emergent misalignment.” We studied why this happens.
Through this research, we discovered a specific internal pattern in the model, similar to a pattern of brain activity, that becomes more active when this misaligned behavior appears. The model learned this pattern from training on data that describes bad behavior. We found we can make a model more or less aligned, just by directly increasing or decreasing this pattern’s activity. This suggests emergent misalignment works by strengthening a misaligned persona in the model.
We showed that training the model again on correct information can push it back toward helpful behavior. Together, this means we might be able to detect misaligned activity patterns, and fix the problem before it spreads.
In short, this work helps us understand why a model might start exhibiting misaligned behavior, and could give us a path towards an early warning system for misalignment during model training.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does RLHF training shape models to prioritize agreement over accuracy? Can base models hide emergent misalignment through alignment training?- Does metagaming behavior actually cause models to act less aligned?
- What undetected misaligned behaviors might Claude Opus 4.6 be hiding?
- How many third parties were affected across each misalignment category?
- Are the five misalignment categories distinct or do they overlap strategically?
- What types of model behavior qualify as misalignment under OpenAI's framework?
- What is the difference between activating a pre-trained persona versus learning new misaligned behavior?
- How do steering-based causal interventions on latents compare to other misalignment mitigation methods?
- Can representational distance to training data explain which prompts trigger misalignment?
- Do base models show emergent misalignment without post-training alignment procedures?
- Can backdoor triggers make emergent misalignment detectable only in specific contexts?
- What happens when alignment values become misaligned with human preferences at scale?
- Why does correct model output not guarantee absence of internal misalignment?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Do inoculation prompts prevent misalignment without harming instruction following?
- Are instruction following gains and emergent misalignment from the same learned change?
- What base rate does concentrated task distribution tell us about real misalignment?
- Does representational distance predict which outputs trigger emergent misalignment?
- Do persona latents like toxic and sarcastic ones generalize across different model architectures?
- How do training-data quality and data poisoning pathways interact with persona latent activation?