跳到正文
北京时间
原文
🚨 AI News | TestingCatalog· @testingcatalog · X·· 23 天前精选AI 评分77
AI 导读

GPT-6 Astra (Max) 以 1797 分登顶 Code Arena: WebDev,领先第 2 名 Claude Fable 5.1 (Max) 35 分、第 3 名 Claude Opus 5 (Max) 1688 分。

推荐理由

榜单数据给出了与 Claude 系列的具体分差和同价位对比,读者可以据此评估新模型的实际编码位置。

正文 · AI 翻译

OpenAI 的 GPT-6 Astra 在 WebDev Arena 上登顶,以 35 分的优势超越了近期发布的 Claude Fable 5.1。

“它还重塑了帕累托前沿,成为 $40/Mtoken 价位上性能最佳的模型,这与最新 Claude 模型的定价一致。”

我们今年年底前能达到 2k 吗?

引用Arena.ai@arena
真实世界的结果已经出炉。Code Arena 上出现了新的 #1 —— GPT-6 Astra (Max)! 它还重塑了帕累托前沿,成为每百万 token 40 美元价位上表现最佳的模型,这与最新的 Claude 模型定价相当。 @OpenAI 的 GPT-6 Astra 在 Code Arena: WebDev 中以 1797 分登顶。这比排名第 2 的 Claude Fable 5.1 (Max)(1762 分)领先了整整 +35 分,比排名第 3 的 Claude Opus 5 (Max)(1688 分)也更高。 相比排名第 13 的 GPT-5.6 Sol (xHigh),这是一次 +180 分的显著提升。 分类级别的投票仍在统计中,但我们已经看到它在以下类别中排名第 1:数据与分析、消费产品、内容创作工具,并在游戏与模拟中排名第 2。 敬请关注其他类别,如品牌与营销、基于参考的设计以及全栈排名。 祝贺 @OpenAI 团队发布这一成果!
原文

Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing. GPT-6 Astra by @OpenAI takes the top spot in Code Arena: WebDev with a score of 1797 pts. This opens up a solid +35pt lead over #2 Claude Fable 5.1 (Max) at 1762 pts and #3 Claude Opus 5 (Max) at 1688 pts. This is a significant improvement from GPT-5.6 Sol (xHigh) at +180 pts, ranked at #13. Category level votes still incoming, but already we see it at #1 in: Data & Analytics, Consumer Product, Content Creation Tools and #2 in Gaming and Simulations. Stay tuned for other categories like Brand & Marketing, Reference-Based Design and Full Stack rankings. Congrats to the @OpenAI team on this release!

在 X 查看被引用的帖子

来源:🚨 AI News | TestingCatalog · x.com

相关事件