The new model tops Anthropic’s internal coding and knowledge-work benchmarks while keeping the same price as its predecessor but the company says it’s intentionally holding back on cybersecurity capability.
Anthropic released Claude Opus 5 on July 24, positioning it as the new default model for coding, agentic tasks, and professional work across its consumer and developer products. The company says the model comes close to the frontier intelligence of Claude Fable 5 at half the price, and it is now the default model on Claude Max and the strongest option on Claude Pro
Pricing is unchanged from the outgoing Opus 4.8: $5 per million input tokens and $25 per million output tokens. Anthropic is also offering a “Fast mode” that runs around 2.5 times the default speed for twice the base price.
What’s Actually New
Anthropic’s own benchmark claims are the centerpiece of the announcement. On the company’s Frontier-Bench v0.1 coding evaluation, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench, Anthropic says the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task.
The company also reports gains outside coding: on the ARC-AGI 3 reasoning test, Opus 5’s score is three times as high as the next-best model, and on Zapier’s AutomationBench, its pass rate is around 1.5× the next-best model for the same cost per task. Anthropic says the model made its largest jumps in organic chemistry and protein-function prediction within its life-sciences evaluation suite.
Early-access partners echoed similar findings. Cognition CEO Scott Wu said Opus 5 approaches Fable-level performance at half the cost within Devin, while Cursor co-founder Sualeh Asif described it as delivering near Fable 5 intelligence at Opus speed and cost. As with all vendor-supplied testimonials, these come from companies with existing commercial relationships with Anthropic and should be read in that context.
All of these figures are Anthropic’s own internal or partner-run evaluations rather than independent third-party benchmarks, a standard caveat for AI model launches the numbers describe controlled test conditions, not necessarily typical real-world usage.
Safety Posture: Capable, But Deliberately Constrained on Cyber
Notably, Anthropic frames Opus 5 as not pushing the capability frontier in one specific area: offensive cybersecurity. The company says it has intentionally avoided training Opus 5 on cyber tasks, and while the model has grown more capable at finding vulnerabilities almost on par with Anthropic’s more restricted Mythos 5 model, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities turning a flaw into a working exploit.
Practically, this means Opus 5’s safety classifiers block binary-based vulnerability scanning, penetration testing, and exploit generation, though Anthropic says those classifiers will intervene around 85% less often than they do for Fable 5. Flagged requests fall back automatically to Opus 4.8 by default in Claude’s consumer apps.
On the alignment side, Anthropic’s internal audit found Opus 5 to be its most aligned model to date, reporting lower rates of deceptive behavior and reduced susceptibility to misuse than Opus 4.8, Sonnet 5, or Fable 5, again, a self-reported metric rather than an externally verified one.
What to Watch
Alongside Opus 5, Anthropic is rolling out two developer-facing betas: mid-conversation tool switching without invalidating prompt caches, and automatic model fallbacks for requests blocked by safety classifiers. The bigger open question is how Opus 5’s real-world coding performance holds up against independent benchmarks and competing frontier models from OpenAI and Google, now that Anthropic has effectively priced it as a mid-tier option punching at near-flagship weight.



