POLICY · REGULATION · #508
Anthropic and OpenAI propose embedding independent safety evaluators inside frontier AI firms
Anthropic CEO Dario Amodei published a proposal to embed third‑party evaluators (e.g., METR, Redwood Research) inside frontier AI companies with unprecedented access to systems, checkpoints, logs, and the ability to publish key findings; OpenAI CEO Sam Altman signaled similar support. Evaluators broadly welcomed the idea but warned that details—what access, contractual controls, publication rights, and legal backing—are unresolved and determine whether such teams would be truly independent or contractors constrained by NDAs and developer control.
KEY POINTS
- Anthropic CEO Dario Amodei published a proposal to embed third‑party evaluators (e.g., METR, Redwood Research) inside frontier AI companies with unprecedented access to systems, checkpoints, logs, and the ability to publish key findings; OpenAI CEO Sam Altman signaled similar support.
- Evaluators broadly welcomed the idea but warned that details—what access, contractual controls, publication rights, and legal backing—are unresolved and determine whether such teams would be truly independent or contractors constrained by NDAs and developer control.
- Embedding truly independent evaluators with deep access could reveal training‑time failures and reduce the risk of models gaming safety tests, but independence depends on contractual and legal safeguards.
WHY IT MATTERS
Embedding truly independent evaluators with deep access could reveal training‑time failures and reduce the risk of models gaming safety tests, but independence depends on contractual and legal safeguards.