Arena.ai· @arena · X·· 15 小时前精选AI 评分68
AI 导读
Arena 公布 Claude Opus 5.5 (High) 以 +12.15% 净提升分数进入 Agent Arena 第二名,仅低于 Claude Fable 5.1 (Max)。其每任务中位价格为 $1.31,比同水平低 64%,成本比 Opus 5 (High) 低 40%、比 Opus 5 (Max) 低 56%;分项信号中 Steerability 排名第一(+14.50%)。
正文 · 原文
ICYMI @AnthropicAI’s Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier.
Opus 5.5 (high) not only improved upon both Opus 5 variants with a higher net improvement score than either, but does so at at 40–56% lower cost:
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release!在 X 查看被引用的帖子
来源:Arena.ai · x.com