Unsloth AI· @UnslothAI · X·· 25 天前精选AI 评分71
AI 导读
Unsloth 宣布通过优化解码和新增 MTP 支持,让 GLM-5.3-Flash 的本地 GGUF 推理提速 1.6 至 3.4 倍,长上下文下最高达 3.3 倍。
推荐理由
原文给出本地运行 GLM-5.3-Flash 的具体加速倍数、显存需求表和现成 GGUF 资源,方法可直接复用。
正文 · 原文
We made GLM-5.3-Flash run 3.3x faster locally!
Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.
Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp.
Guide: https://unsloth.ai/docs/models/glm-5.3-flash#faster-inference-and-mtp-support
GGUF: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai在 X 查看被引用的帖子
来源:Unsloth AI · x.com