Arena 宣布 Anthropic 的 Claude Opus 5.5 已进入 Agent Arena,用户可通过投票驱动排行榜。Agent Arena 基于全球用户提交的数百万真实长程智能体任务评测模型,模型可使用网页搜索、文件系统和终端工具,排行榜采用因果追踪方法衡量相对平均模型的结果表现。
Arena 官宣 Claude Opus 5.5 上架 Agent Arena 和 Battle Mode,读者可以参与投票影响排行榜结果。
Claude Opus 5.5 by @AnthropicAI is now in the Agent Arena!
Your votes drive the @arena leaderboards, head over and bring your toughest prompts.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
In addition to Agent Arena, @claudeai Opus 5.5 is in Battle Mode for: WebDev, Text, Vision, and Document.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.在 X 查看被引用的帖子
来源:Arena.ai · x.com