Anthropic released Claude Opus 5 on Friday, and it immediately became the day's dominant story, topping Hacker News with more than 1,400 points and taking the number-one slot on the Artificial Analysis intelligence leaderboard. The headline claim is a favorable point on the cost-capability curve: Anthropic says Opus 5 comes close to the frontier intelligence of its flagship Fable 5 while costing half as much, and it ships at the same price as its predecessor, Opus 4.8 — five dollars per million input tokens and twenty-five per million output. It becomes the default model on Claude Max and the strongest option on Claude Pro.
The benchmark story is broad rather than narrow. On the internal Frontier-Bench software-engineering suite Opus 5 is the new state of the art, more than doubling Opus 4.8's score at a lower cost per task; on CursorBench it lands within half a percentage point of Fable 5's peak at half the cost. On ARC-AGI 3, a test of solving genuinely novel problems, its score is roughly three times the next-best model. It leads on the OSWorld computer-use benchmark at any given cost, and on Zapier's business-automation test its pass rate is about one and a half times the next-best system. Anthropic also reports across-the-board gains on life-sciences evaluations, with the largest jumps in organic chemistry and protein variant-effect prediction. The company stresses behavior as much as scores: Opus 5 is described as more willing to verify its own work and iterate, with partners at Cognition, Cursor, Zapier and Lovable emphasizing lower run-to-run variance over peak numbers.
The alignment and safety framing is unusually detailed. Anthropic's automated behavioral audit rates Opus 5 its most aligned model to date, with the lowest measured rate of misaligned behavior of any recent Claude, the lowest deceptive-behavior rate, and the least susceptibility to being tricked into misuse. On dangerous capabilities the company is careful to say Opus 5 does not advance the frontier: it remains behind the higher-tier Mythos 5 in both biology and offensive cybersecurity. On an internal fuzzing benchmark it identifies software vulnerabilities nearly as well as Mythos 5 but is far weaker at turning them into working exploits — the step Anthropic treats as the real threat. Its cyber classifiers are tuned to be about eighty-five percent less restrictive than Fable 5's, permitting source-code vulnerability discovery while blocking binary-based scanning, penetration testing and exploit generation, with flagged requests falling back to Opus 4.8. Alongside the model, Anthropic shipped two API betas: mid-conversation tool changes that don't invalidate the prompt cache, and automatic fallbacks that reroute safety-flagged requests to another model rather than refusing outright.
- Anthropic frames Opus 5 as reaching Fable 5-class intelligence at half the cost, and the new default on Claude Max.
- Hacker News made it the runaway top story of the day (1,482 points, 814 comments); a second thread flagged Opus 5 taking the #1 slot on the Artificial Analysis Intelligence leaderboard.
- TechCrunch's read: cheaper and less cyber-restricted than Fable, likely making it the default choice for most workloads.
- Early-access partners (Cognition, Cursor, Zapier, Lovable) emphasized reliability and lower run-to-run variance over peak scores.