跳到正文
北京时间
原文
Arena.ai· @arena · X·· 14 小时前精选AI 评分69
AI 导读

Arena 宣布 Anthropic 的 Claude Sonnet 5.5 已进入 Agent Arena,并开放投票。Agent Arena 基于全球用户数百万个真实的长程智能体任务评测模型,模型可使用 web search、filesystem 和 terminal 工具完成复杂工作流,榜单用因果追踪方法衡量模型相对平均模型的结果表现。

推荐理由

Arena 介绍了其 Agent Arena 的评测方式,读者可以据此理解 Claude Sonnet 5.5 在真实智能体任务上的衡量口径。

正文 · 原文

Claude Sonnet 5.5 by @AnthropicAI is now in the Agent Arena!

Your votes drive the @arena leaderboards, head over and bring your toughest prompts.

In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.

In addition to Agent Arena, @claudeai Sonnet 5.5 is in Battle Mode for: WebDev, Text, Vision, and Document.

引用Claude@claudeai
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
在 X 查看被引用的帖子

来源:Arena.ai · x.com