跳到正文
北京时间
原文
Arena.ai· @arena · X·· 3 小时前精选AI 评分69
AI 导读

Arena 发布 Agent Arena 最新榜单,Anthropic 的 Claude Sonnet 5.5 以 +12.5% 净提升得分排名第 3,单任务中位成本 $2.74,比排名第 2 的 Claude Opus 5.5($1.58)高约 73%,且 Opus 5.5 得分更高,因此 Sonnet 5.5 未进入 Agent Arena 的 Pareto 前沿。据引用内容,Sonnet 5.5 在 Chat 类目以 +15.6% 排名第 1,Anthropic 模型包揽 Agent Arena 前三名。

推荐理由

原文基于 Arena 榜单数据指出 Claude Sonnet 5.5 得分高但成本更高、未进 Pareto 前沿,读者可以据此比较成本与性能的取舍。

正文 · AI 翻译

Claude Sonnet 5.5 由 @AnthropicAI 发布,刚刚在 Agent Arena 中排名第 #3。该模型每项任务的中位成本为 $2.74,净提升得分为 +12.5%。

Claude Sonnet 5.5 提供了顶级性能,但代价高昂:#2 的 Claude Opus 5.5 每项任务成本为 $1.58,同时取得了更高的得分。这一权衡使 Sonnet 5.5 刚好落在 Agent Arena 帕累托前沿之外。

引用Arena.ai@arena
激动人心的消息:@AnthropicAI 的 Claude Sonnet 5.5 (Max) 在 Agent Arena 中首次亮相便位列第 3,净提升达 +12.5%! 此次发布相比 Claude Sonnet 5 (High) 提升了 8.1 个百分点,后者以 +4.4% 的净提升位列第 13。按类别划分,Claude Sonnet 5.5 在 Chat 中夺得第 1 名(+15.6%),高于 Fable 5.1(+11.49%)和 Opus 5.5(+10.29%)。 这一表现伴随着更高的成本:Claude Sonnet 5.5 (Max) 的每任务中位成本为 $2.74,比排名第 2 的 Claude Opus 5.5 (High) 的 $1.58 高出约 73%。 @AnthropicAI 的模型如今占据了 Agent Arena 的前三名。恭喜团队!
原文

Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!

在 X 查看被引用的帖子

来源:Arena.ai · x.com