DeepSeek-V4.1-Flash (Max) 进入 Agent Arena 开源模型第 3 名,净提升 +4.87%,每任务中位成本 $0.07,重塑 Pareto 前沿。其成本比第 2 名 Hy4 preview 低 68%、成绩仅差 0.09 个百分点;总榜排名第 12,Confirmed Success 信号排名 第 4(+13.75%)。
原文给出开源模型成本与成绩的完整对比,读者可以据此评估 DeepSeek-V4.1-Flash 在性价比上的位置。
Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier.
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median cost per task. Its +4.87% net improvement is within 0.09 percentage points of Hy4 preview (ranked #2) at 68% lower cost, and within 1.52 percentage points of Kimi K3 (Max) (ranked #1) at 91% lower cost.
- Kimi K3 (Max): +6.39% | $0.77/task
- Hy4 preview: +4.96% | $0.22/task
- DeepSeek-V4.1-Flash (Max): +4.87% | $0.07/task
See the full Pareto Frontier below.
DeepSeek-V4.1-Flash (Max) is ranked #12 overall, and by signal landed #4 Confirmed Success with +13.75%!
Congrats to the @deepseek_ai team on this release!
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6在 X 查看被引用的帖子
来源:Arena.ai · x.com