跳到正文
北京时间
原文
Cognition 模型 / Devin 博客(网页)·· 2026-04-07精选AI 评分65

Cognition 发布软件工程智能体模型 SWE-1.6,上线 Windsurf

Introducing SWE 1.6: Improving Model UX

AI 导读

Cognition 发布面向软件工程智能体的新模型 SWE-1.6,已在 Windsurf 全面开放,未来 3 个月免费。模型在 SWE-Bench Pro 上与 SWE-1.6 Preview 表现相当,同时显著减少过度思考、循环推理和依赖终端等行为,更多使用并行工具调用;训练中引入长度惩罚,响应长度增长更慢且任务解决率保持稳定。

推荐理由

官方说明了用长度惩罚等训练手段改善模型 UX 的具体做法和效果,对关注智能体行为质量的人有可参考的细节。

正文

We’re releasing SWE-1.6, our latest model built for software engineering agents, and we’re making it generally available in Windsurf. SWE-1.6 is optimized for both intelligence and model UX. Moreover, it is industry-leading in both speed (up to 950 tok/s) and cost (free tier for the next 3 months).

Last month, we released SWE-1.6 Preview, which improved on SWE-Bench Pro by more than 10% compared to our previous model SWE-1.5 while being post-trained on the same pre-trained model. SWE-1.6 was post-trained from scratch to jointly optimize for user experience and making the model feel smoother to use in addition to raw intelligence.

Introducing SWE 1.6: Improving Model UX

While SWE-1.6 achieves comparable performance to the Preview model on benchmarks like SWE-Bench Pro, we’re most excited about its dramatic improvement in what we call “model UX”. As we observed in our earlier post, the preview checkpoint exhibited several behavioral issues that added friction for our users. These included:

  • Overthinking for simple problems, taking more turns than necessary for simple tasks.
  • Calling tools sequentially rather than in parallel
  • Preferring shell commands rather than its own tools
  • Exhibiting “looping behavior”, getting caught in a circle of identical reasoning

Many of these axes aren’t measured by traditional benchmarks but significantly affect the infamous “vibes” users express when trying the model.

Introducing SWE 1.6: Improving Model UX

We were able to significantly reduce the frequency of such behaviors in SWE 1.6. The model now uses parallel tool calls more often, loops far less and relies more on its tools than the terminal. This leads to more efficient trajectories and a smoother user experience: the model obtains context much faster and requires less input from the user.

In the example below, when asked a question about the PyTorch codebase, SWE-1.6 uses parallel tool calls far more than the preview and answers the question faster.

视频 · 前往原文观看
SWE-1.6 uses parallel tool calls much more than the preview.

One contributing factor to this improvement was the introduction of a length penalty into training, which discourages unnecessarily long trajectories. This directly reduces overthinking and looping, while implicitly encouraging more efficient behaviors like parallel tool use. During training, we observed the model response length growing much more slowly than before while maintaining its intelligence and coding ability. The below ablation shows that task solve rate stays similar while assistant turns stays flat.

Introducing SWE 1.6: Improving Model UX

We also were able to significantly reduce occurrences of the model relying on the terminal and other improper tool use cases throughout training, avoiding cases where Windsurf users have to manually accept commands instead than letting the agent work continuously.

Introducing SWE 1.6: Improving Model UX

Try SWE-1.6

SWE 1.6 is available for everyone today in Windsurf, and will be free for the next 3 months. We have partnered with Fireworks to offer the free version at 200 tok/s. We have also partnered with Cerebras to offer a faster version of the model for our paying users at 950 tok/s, delivering the same intelligence with unmatched speed and cost.

视频 · 前往原文观看
SWE-1.6 Fast delivers output at 950 tok/s.

Try them both in Windsurf today!

来源:Cognition 模型 / Devin 博客(网页) · cognition.com