跳到正文
北京时间
原文
Simon Willison 博客· Simon Willison·· 2026-04-23精选AI 评分71

Qwen3.6-27B:27B 稠密模型实现旗舰级编程能力

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

AI 导读

Qwen 发布 Qwen3.6-27B 开源模型,这款 27B 参数的稠密模型在编码基准测试中超越了上一代 397B 总参数/17B 激活参数的 MoE 旗舰模型 Qwen3.5-397B-A17B。模型体积从 807GB 大幅缩减至 55.6GB,16.8GB 的量化版本即可在本地运行,生成速度约 25 tokens/s。实测显示其能生成复杂的 SVG 图形代码,展现出旗舰级的编程能力。

推荐理由

通义千问 3.6 27B 用 55GB 瘦身换来了超 397B MoE 的编码能力,Simon 亲手跑出的 SVG 效果惊艳,2026 年最值得本地部署的开源小模型。

正文 · AI 翻译

Simon Willison 的博客

2026年4月22日 - 链接博客

Qwen3.6-27B:27B 稠密模型中的旗舰级编码能力(via)通义千问对其最新开源权重模型做出了重大宣称:

Qwen3.6-27B 提供了旗舰级的智能体编码性能,在所有主要编码基准测试中超越了上一代开源旗舰模型 Qwen3.5-397B-A17B(总计 397B / 17B 激活 MoE)。

在 Hugging Face 上,Qwen3.5-397B-A17B 大小为 807GB,而这款新的 Qwen3.6-27B 仅为 55.6GB。

我使用 16.8GB 的 Unsloth Qwen3.6-27B-GGUF:Q4_K_M 量化版本和 llama-server 进行了测试,参考了 Hacker News 上 benob 的这份配方,首先通过 `brew install llama.cpp` 安装了 llama-server:

llama-server \
    -hf unsloth/Qwen3.6-27B-GGUF:Q4_K_M \
    --no-mmproj \
    --fit on \
    -np 1 \
    -c 65536 \
    --cache-ram 4096 -ctxcp 2 \
    --jinja \
    --temp 0.6 \
    --top-p 0.95 \
    --top-k 20 \
    --min-p 0.0 \
    --presence-penalty 0.0 \
    --repeat-penalty 1.0 \
    --reasoning on \
    --chat-template-kwargs '{"preserve_thinking": true}'

首次运行时,它将约 17GB 的模型保存到了 `~/.cache/huggingface/hub/models--unsloth--Qwen3.6-27B-GGUF`。

以下是“生成一只骑自行车的鹈鹕的 SVG”的对话记录。对于一个 16.8GB 的本地模型来说,这是一个非常出色的结果:

llama-server 报告的性能数据:

  • 读取:20 tokens,0.4 秒,54.32 tokens/s
  • 生成:4,444 tokens,2 分 53 秒,25.57 tokens/s

为了更全面地展示,这里是“生成一只北弗吉尼亚负鼠骑电动滑板车的 SVG”(之前用 GLM-5.1 运行过):

那次生成了 6,575 tokens,耗时 4 分 25 秒,速度为 24.74 t/s。

2026年4月22日

来源:Simon Willison 博客 · simonwillison.net