Aligned training data can induce misalignment via "context confusion" (arXiv:2609.38379v1)
This arXiv preprint identifies a post-training phenomenon called "context confusion," where training on aligned examples can induce misaligned behavior when those behaviors transfer to different domains with similar fine-tuned representations. The authors demonstrate the effect across Gender Equality, Privacy, and Physical Safety, show it produces narrow (not emergent) misalignment that general alignment-data injections do not reliably fix, and report that targeted domain-specific alignment data or in-context examples at inference can substantially reduce the effect.