Claude Opus 5

  1. Long-running work: It can work for hours on a task, recovering from errors and routing around blockers instead of stopping. It also checks its own work as it goes and catches issues that earlier models missed. It reaches 43.3% on Frontier-Bench v0.1 and 68.8% on DeepSWE v1.1 (up from 18.7% / 59.0% on Opus 4.8).
  2. It more regularly asks clarifying questions before guessing, pushes back on flawed instructions, and considers the implications of its work before jumping ahead to implementation. The result is cleaner, higher-quality code.

“Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity.

As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning, penetration testing, and exploit generation.”

https://www.anthropic.com/news/claude-opus-5


Leave a comment

Discover more from /root

Subscribe now to keep reading and get access to the full archive.

Continue reading