RELEASE · MODELS · #397
Anthropic releases Claude Opus 5 — lower-cost model claiming near‑frontier performance
Anthropic announced Claude Opus 5, a new model positioned as a cost‑efficient successor to Opus 4.8 and the new default on Claude Max (and the strongest on Claude Pro). Anthropic says Opus 5 matches or exceeds prior models on many coding, knowledge‑work and scientific benchmarks (Frontier‑Bench, GDPval‑AA, CursorBench, ARC‑AGI, Zapier AutomationBench, OSWorld) at lower cost per task while remaining behind Mythos 5 on security and biology frontier tasks; the company also reports improved alignment and safety in pre‑deployment audits and links a System Card for more details.
KEY POINTS
- Anthropic announced Claude Opus 5, a new model positioned as a cost‑efficient successor to Opus 4.8 and the new default on Claude Max (and the strongest on Claude Pro).
- Anthropic says Opus 5 matches or exceeds prior models on many coding, knowledge‑work and scientific benchmarks (Frontier‑Bench, GDPval‑AA, CursorBench, ARC‑AGI, Zapier AutomationBench, OSWorld) at lower cost per task while remaining behind Mythos 5 on security and biology frontier tasks; the company also reports improved alignment and safety in pre‑deployment audits and links a System Card for more details.
- A model that claims near‑frontier performance at substantially lower cost — plus improved alignment and broad domain gains — could shift enterprise and developer adoption and affects competitive positioning among leading large models.
WHY IT MATTERS
A model that claims near‑frontier performance at substantially lower cost — plus improved alignment and broad domain gains — could shift enterprise and developer adoption and affects competitive positioning among leading large models.
SOURCES & TIMELINE
4Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA , Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the…
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar review…
Today, we are introducing the Life Sciences Verification Program (LSVP), which gives life science professionals access to our Mythos, Opus, and Sonnet models with a refined set of safeguards more permissive for biology-related work. We have already onboarded dozens of organizations through an early-access program, and are now opening applications to the broader life science community ( apply here ). The program is la…
In a twist that captures the strange new state of AI security, independent security researchers have used Anthropic’s Claude to break into OpenAI, exposing cracks in the ChatGPT-maker’s defenses, The Wall Street Journal reported on Thursday evening. A three-person security team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program. Hacktron reported its findings to OpenAI, which gave …
A three-person team of researchers used a corrupted image file and forum software to hack into OpenAI. A team of three independent security researchers at Hacktron says it took less than 72 hours for them to hack into OpenAI employee accounts using Anthropic’s Claude Opus 4.8 and 5, the Wall Street Journal reports. They were able to access OpenAI’s GitHub repository, called “Monorepo,” which reportedly contains “Ope…
Security researchers used Anthropic's Claude to break into OpenAI's internal systems and code repositories. The attack ran through OpenAI's community forum, where chained vulnerabilities gave the researchers access to employee ChatGPT and Codex accounts. The entire exploit took just 72 hours to develop, showing how drastically AI models are cutting the time and skill needed for sophisticated cyberattacks. Three se…