Claude Opus 5 released: approaching Fable 5 cutting-edge intelligence at half the price, surpassing 4.8 in scientific evaluation

Anthropic released Claude Opus 5 on July 24: At the same price as Opus 4.8, it refreshes SOTA on multiple benchmarks such as Frontier-Bench and CursorBench. The ARC-AGI 3 score is 3 times that of the sub-optimal model, calling it the "highest alignment and lowest deception in history" model.

Anthropic dropped a carefully calculated chess piece on July 24: Claude Opus 5 was officially released - the pricing is exactly the same as Opus 4.8, but the performance is "close to the flagship at half the price" and penetrates into the blank areas of its high-end and mid-range product lines. The official positioning is very straightforward: this is a cutting-edge model that can be used daily, and it has naturally become the default model of Claude Max and the most powerful model on Claude Pro.

Benchmark results: SOTA vs. "3x the next best"

Opus 5’s hard data leaves a “crushing” footnote on almost every key scale:

  • Frontier-Bench v0.1: Surpasses all other models and achieves more than twice the performance of Opus 4.8 at a lower cost per task
  • CursorBench 3.2: The highest effort is only 0.5% different from the Fable 5 peak, and the cost is half of it
  • ARC-AGI 3: Scores 3x higher than the next best model
  • Zapier AutomationBench: At the same single task cost, the pass rate is approximately 1.5 times that of the suboptimal model
  • OSWorld 2.0: Beats Fable 5 best result at over one-third the cost

The scientific evaluation is even better than 4.8: all life science evaluations are superior, the organic chemistry internal score is 10.2 percentage points higher than 4.8, and the protein-related tasks are 7.7 percentage points higher.

Pricing and new engineering capabilities

Opus 5 is available on all platforms starting today, priced at $5/million input tokens, $25/million output tokens (same as Opus 4.8), and provides Fast mode (about 2.5 times faster, 2 times the price). Two simultaneous betas are worth the attention of the engineering team: Tool changes during the session (prompt cache will no longer be invalidated, the cost structure of long sessions will be significantly improved) and API automatic rollback (an official attempt at downgrade fault tolerance).

At the level of alignment, Anthropic made a rare statement - Opus 5 is its "most aligned and least deceptive model ever", with an alignment score of 2.3. In line with the "Jailbreak Severity Score" industry framework promoted on June 30, Anthropic is marketing security capabilities as a selling point alongside performance.

From an industry perspective, the profound meaning of the Opus 5 card position is that while Fable 5 (cutting-edge flagship) and Sonnet 5 (large-scale main force) occupy the two ends respectively, Opus 5 fills the middle range with the "half-price flagship". This is actually a direct mirror response to the three-tier pricing of OpenAI GPT-5.6 (Sol/Terra/Luna). For developers, manufacturers on both sides of this "cost-capability" balance are handing over the choice to the same variable - the dollar cost per unit of intelligence.

Several directions worth tracking in the future:

  1. ARC-AGI 3 evaluation methodology: After OpenAI disclosed "two settings triple the score" on July 29, the horizontal comparability of each company's ARC-AGI 3 numbers needs to be re-examined.
  2. The true cost impact of mid-session tool changes: prompt cache retention rate is the core cost variable for long Agent tasks.
  3. Penetration of Opus 5 in scientific computing scenarios: Can the leadership in life science evaluations be translated into actual purchases by medical and biopharmaceutical customers.
  4. Fable 5's follow-up actions: After the flagship is "approached by half price", Anthropic's next-generation flagship upgrade rhythm will be forced to advance.
Copyright: Content sourced from Anthropic Official News . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...