跳到正文
北京时间
原文
Arena.ai· @arena · X·· 14 小时前精选AI 评分68
AI 导读

Arena 公布 Claude Opus 5.5 (High) 以 +12.15% 净提升分数进入 Agent Arena 第二名,仅低于 Claude Fable 5.1 (Max)。其每任务中位价格为 $1.31,比同水平低 64%,成本比 Opus 5 (High) 低 40%、比 Opus 5 (Max) 低 56%;分项信号中 Steerability 排名第一(+14.50%)。

正文 · AI 翻译

如果你错过了:@AnthropicAI 的 Opus 5.5(High)在 Agent Arena 中排名第 2,并重塑了帕累托前沿。

Opus 5.5(high)不仅比两个 Opus 5 变体都有所提升,净改进分数高于两者,而且成本还降低了 40–56%:

引用Arena.ai@arena
来自 @AnthropicAI 的 Claude Opus 5.5 (High) 刚刚进入 Agent Arena,排名第 #2,净提升分数为 +12.15%。只有 Fable 5.1 (Max) 排名更高。然而,Opus 5.5 (High) 每任务中位价格为 $1.31,成本降低了 64%,从而推进了帕累托前沿。 Opus 5.5 (High) 的净提升分数高于之前两个 Opus 5 变体,同时成本比 Opus 5 (High) 低 40%,比 Opus 5 (Max) 低 56%。 按信号划分,Opus 5.5 (High) 排名: - #1 可操控性 (+14.50%) - #2 确认成功 (+15.50%) - #3 表扬与投诉对比 (+19.80%) - #4 Bash 恢复 (+10.64%) 祝贺 @AnthropicAI 又发布了一款前沿模型!
原文

Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release!

在 X 查看被引用的帖子

来源:Arena.ai · x.com