跳到正文
北京时间
原文
Unsloth AI· @UnslothAI · X·· 2026-08-27精选AI 评分72
AI 导读

Unsloth 发布 GLM-5.3-Flash(ox-alpha)的 GGUF 量化版本,可在 128GB RAM 设备上以 3-bit 运行。该模型为 Z.ai 的 320B-A18B 开源多模态模型,MIT License 发布,原文称其在 DeepSWE、编码和智能体基准上媲美 Claude Opus 4.8。

推荐理由

原文给出了各量化档位的内存需求和精度保留数据,读者可据此判断在 128GB 设备上本地运行该模型的可行性。

正文 · 原文

GLM-5.3-Flash can now be run locally! ✨

Run 3-bit on 128GB RAM via Unsloth GGUF.

GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.

Guide: https://unsloth.ai/docs/models/glm-5.3-flash
GGUF: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

引用Z.ai@Zai_org
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai
在 X 查看被引用的帖子

来源:Unsloth AI · x.com