Arena 公布 Claude Opus 5.5 (High) 进入 Agent Arena 排名第 2,净改进分 +12.15%,仅次于 Fable 5.1 (Max)。
榜单给出了 Claude Opus 5.5 (High) 在 Agent Arena 的排名、价格与分项信号,读者可据此权衡它与 Opus 5 各档的成本与能力变化。
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier.
Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max).
By signal, Opus 5.5 (High) ranks:
- #1 Steerability (+14.50%)
- #2 Confirmed Success (+15.50%)
- #3 Praise vs Complaint (+19.80%)
- #4 Bash Recovery (+10.64%)
Congrats to the @AnthropicAI on another frontier model release!
Big news: Claude Opus 5.5 (Max) by @AnthropicAI just topped #1 in Code Arena: WebDev with 1818 pts and reshapes the Pareto frontier! This is a solid +26pt lead ahead of the next best model, GPT-6 Astra (Max) and a huge +126pt improvement over previous Opus 5 (Max) at 1692. Across categories, we can see that Opus 5.5 (Max) lands in the top spots across domains for: - #1 Brand and Marketing, Reference-Based Design, Data & Analytics, Simulations and Gaming! - #2 Consumer Product Stay tuned for more domain and categorical insights to land as more votes come in. Congrats to @AnthropicAI on the release!在 X 查看被引用的帖子
来源:Arena.ai · x.com