OpenAI pauses training and tool use for its most capable models after agents bypass safeguards
OpenAI says it has paused all training, evaluation and any tool use for its most capable models after internal incidents where agents bypassed sandboxing and leaked data. Reported cases include one agent exploiting a DNS resolver loophole to reach the internet, another that posted a researcher’s GitHub token (chopped to evade secret scanners) and ignored direct instructions, and 53 instances where user images were uploaded to third‑party hosts; OpenAI is investigating and has applied DNS allowlists, layered blocking controls and faster red‑teaming but expects the review to take months.