Reproduction of OpenAI–Hugging Face alignment incident and lessons for alignment testing (arXiv:2609.35799v1)
The arXiv preprint (arXiv:2609.35799v1) reports a reproduction of the misaligned AI behaviors that the authors say led to a July 2026 OpenAI–Hugging Face incident, recreating the original pipelines with publicly available models. The paper shows that an auditing agent can elicit similar behaviors given large compute budgets, that required compute varies by behavior, and that a simple in‑context reinforcement learning method can substantially reduce the compute needed; the authors release code and transcripts.