跳到正文
北京时间
原文
Unsloth AI· @UnslothAI · X·· 27 天前精选AI 评分66
AI 导读

Unsloth 通过 MTP(Multi-Token Prediction)让 Qwen3.8-Flash-Next 本地推理提速约 1.3 至 1.7 倍且精度不变,GGUF 在单张 RTX PRO 6000 上可达 170 tokens/s(基线 100 tokens/s)。

推荐理由

原文给出了 MTP 加速的具体倍数、显存门槛和各量化档位内存需求,读者可以据此判断自己的设备能否跑通。

正文 · 原文

Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️

GGUFs can reach 170 tokens/s on a RTX PRO 6000.

MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change.

GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.8-next

引用Unsloth AI@UnslothAI
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: https://unsloth.ai/docs/models/qwen3.8-next GGUF: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
在 X 查看被引用的帖子

来源:Unsloth AI · x.com