Theme of inquiry

How do alignment methods inadvertently affect model behavior and generalization?

A question within its area, explored through 1 line of inquiry below — each a family of specific questions the research asks.


Can base models hide emergent misalignment through alignment training?

117 specific questions

See all 117 questions in this line of inquiry