РЕЛИЗ · MODELS · #397
Anthropic выпускает Claude Opus 5 — недорогая модель с близкой к фронтирной мощностью
Anthropic объявила Claude Opus 5 — новую модель, позиционируемую как более экономичный преемник Opus 4.8 и новый стандарт по умолчанию в Claude Max (а также сильнейшую в Claude Pro). По данным Anthropic, Opus 5 достигла или превзошла предыдущие модели на ряде бенчмарков для программирования, задач знания и науки (Frontier‑Bench, GDPval‑AA, CursorBench, ARC‑AGI, Zapier AutomationBench, OSWorld) при меньшей стоимости выполнения, но отстаёт от Mythos 5 в задачах по кибербезопасности и биологии; компания также указывает на улучшенную выравненность и безопасность и публикует System Card.
КЛЮЧЕВЫЕ ТЕЗИСЫ
- Anthropic объявила Claude Opus 5 — новую модель, позиционируемую как более экономичный преемник Opus 4.8 и новый стандарт по умолчанию в Claude Max (а также сильнейшую в Claude Pro).
- По данным Anthropic, Opus 5 достигла или превзошла предыдущие модели на ряде бенчмарков для программирования, задач знания и науки (Frontier‑Bench, GDPval‑AA, CursorBench, ARC‑AGI, Zapier AutomationBench, OSWorld) при меньшей стоимости выполнения, но отстаёт от Mythos 5 в задачах по кибербезопасности и биологии; компания также указывает на улучшенную выравненность и безопасность и публикует System Card.
- Модель, претендующая на близкую к фронтирной эффективность при значительно меньшей стоимости и с улучшенной выравненностью, может изменить выбор корпоративных и разработческих решений и повлиять на конкурентную динамику рынка больших моделей.
ПОЧЕМУ ЭТО ВАЖНО
Модель, претендующая на близкую к фронтирной эффективность при значительно меньшей стоимости и с улучшенной выравненностью, может изменить выбор корпоративных и разработческих решений и повлиять на конкурентную динамику рынка больших моделей.
ИСТОЧНИКИ И ХРОНОЛОГИЯ
4Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA , Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the…
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar review…
Today, we are introducing the Life Sciences Verification Program (LSVP), which gives life science professionals access to our Mythos, Opus, and Sonnet models with a refined set of safeguards more permissive for biology-related work. We have already onboarded dozens of organizations through an early-access program, and are now opening applications to the broader life science community ( apply here ). The program is la…
In a twist that captures the strange new state of AI security, independent security researchers have used Anthropic’s Claude to break into OpenAI, exposing cracks in the ChatGPT-maker’s defenses, The Wall Street Journal reported on Thursday evening. A three-person security team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program. Hacktron reported its findings to OpenAI, which gave …
A three-person team of researchers used a corrupted image file and forum software to hack into OpenAI. A team of three independent security researchers at Hacktron says it took less than 72 hours for them to hack into OpenAI employee accounts using Anthropic’s Claude Opus 4.8 and 5, the Wall Street Journal reports. They were able to access OpenAI’s GitHub repository, called “Monorepo,” which reportedly contains “Ope…
Security researchers used Anthropic's Claude to break into OpenAI's internal systems and code repositories. The attack ran through OpenAI's community forum, where chained vulnerabilities gave the researchers access to employee ChatGPT and Codex accounts. The entire exploit took just 72 hours to develop, showing how drastically AI models are cutting the time and skill needed for sophisticated cyberattacks. Three se…