Arena.ai· @arena · X·· 4 小时前精选AI 评分69
AI 导读
Arena 发布 Agent Arena 最新榜单,Anthropic 的 Claude Sonnet 5.5 以 +12.5% 净提升得分排名第 3,单任务中位成本 $2.74,比排名第 2 的 Claude Opus 5.5($1.58)高约 73%,且 Opus 5.5 得分更高,因此 Sonnet 5.5 未进入 Agent Arena 的 Pareto 前沿。据引用内容,Sonnet 5.5 在 Chat 类目以 +15.6% 排名第 1,Anthropic 模型包揽 Agent Arena 前三名。
推荐理由
原文基于 Arena 榜单数据指出 Claude Sonnet 5.5 得分高但成本更高、未进 Pareto 前沿,读者可以据此比较成本与性能的取舍。
正文 · 原文
Claude Sonnet 5.5 by @AnthropicAI just landed at #3 in the Agent Arena. This model has a median cost per task of $2.74, and a +12.5% net improvement score.
Claude Sonnet 5.5 delivers top-tier performance, but at a cost premium: #2 Claude Opus 5.5 costs $1.58 per task while achieving a higher score. That tradeoff keeps Sonnet 5.5 just off the Agent Arena Pareto frontier.
Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!在 X 查看被引用的帖子
来源:Arena.ai · x.com