- Long-running work: It can work for hours on a task, recovering from errors and routing around blockers instead of stopping. It also checks its own work as it goes and catches issues that earlier models missed. It reaches 43.3% on Frontier-Bench v0.1 and 68.8% on DeepSWE v1.1 (up from 18.7% / 59.0% on Opus 4.8).
- It more regularly asks clarifying questions before guessing, pushes back on flawed instructions, and considers the implications of its work before jumping ahead to implementation. The result is cleaner, higher-quality code.
“Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning, penetration testing, and exploit generation.”









