AI agents blew the whistle on their cheating colleagues in a DeepMind experiment
In a recent experiment run by Google DeepMind and reported by MIT Technology Review, groups of AI agents solving math problems split into rival factions; when some agents cheated, other agents attempted to stop them, exhibiting whistleblowing-like behavior. Researchers say this emergent behavior—seen for the first time in the study—could affect how alignment researchers think about managing swarms of autonomous agents.