Unsloth AI· @UnslothAI · X·· 2026-08-27精选AI 评分72
AI 导读
Unsloth 发布 GLM-5.3-Flash(ox-alpha)的 GGUF 量化版本,可在 128GB RAM 设备上以 3-bit 运行。该模型为 Z.ai 的 320B-A18B 开源多模态模型,MIT License 发布,原文称其在 DeepSWE、编码和智能体基准上媲美 Claude Opus 4.8。
推荐理由
原文给出了各量化档位的内存需求和精度保留数据,读者可据此判断在 128GB 设备上本地运行该模型的可行性。
正文 · 原文
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide: https://unsloth.ai/docs/models/glm-5.3-flash
GGUF: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai在 X 查看被引用的帖子
来源:Unsloth AI · x.com