跳到正文
北京时间
原文
StepFun· @StepFun_ai · X·· 2026-05-31精选AI 评分80
AI 导读

阶跃星辰发布了Step 3.7 Flash,这是一款198B参数的视觉模型,旨在DGX Spark等桌面设备上运行。用户实测表明,128GB统一内存是运行门槛,模型占用约104GB。部署无需官方专用llama.cpp分支,主线版本即可。在上下文长度上存在权衡:启用视觉功能时,基于q8 KV cache的64K为上限;若要使用最高256K上下文,则需禁用视觉并切换至q4 KV cache,此时模型与缓存共占约114GB内存。该模型是推理模型,思考过程可能消耗大量max_tokens,需注意设置。

推荐理由

把 198B 的视觉模型塞进一台桌面盒子,还跑通了,这本身就是个小里程碑。更关键的是,这篇实战直接帮你绕开了三个大坑,省下的三小时够你喝杯咖啡慢慢试了。

正文 · AI 翻译

一个198B参数的视觉模型,运行在桌面上的一个小机箱里。这就是我们打造 Step 3.7 Flash 的目的。

精彩的拆解分析 @sudoingX — 为大家省去了几个小时的困惑时间 🎉

引用Sudo su@sudoingX
我现在正在一台 dgx spark 上运行 stepfun 新的 step 3.7 flash。 198b 的视觉模型,跑在一台就摆在桌上的机器上。以下是如何帮你省下大约 3 小时抓耳挠腮的加载时间,因为我已经替你抓过了。 官方 README 告诉你需要 stepfun 自己的 llama.cpp fork。你不需要,主线 ggml-org 就能跑得好好的,视觉功能什么的都行,64k。别花一个小时去构建一个 fork,结果发现不用它模型也能加载。 下面这个才是真正会吃掉你一整晚的:模型占了你约 121gb 统一内存池中的 104gb,而 spark 没有 swap。请求的上下文太大,它不会干净地崩溃,而是会静默地颠簸,内核把模型页换出、再从磁盘反复读回,循环往复,而你只能盯着“loading”发呆。 判断的信号是 read_bytes。如果它涨过 104gb 的模型大小还继续涨,同时进程内存钉在 99% 且日志毫无进展,那就是卡死了。杀掉它,它不会自己恢复。 在 q8 kv cache 上加载视觉投影器时,64k 就是 128gb 机器上的上限,超过它就会在 clip loader 里颠簸,模型加上 KV 加上视觉缓冲区全都在争抢你剩下的那约 17gb。 想要完整的 256k 上下文?去掉视觉,换成 q4 kv cache,这样就能装下,121gb 里用 114gb。这就是他们 README 没有明说的取舍:大上下文是真的,只不过得用 q4 cache,而且仅限文本。 而且它是个推理模型。问它点东西,如果你得到空白回复,不是你把它弄坏了,是你把 max_tokens 设得太低,它把整个预算都花在思考上了。答案就在 reasoning_content 里。给它留点空间。 这就是那 3 小时。确切可用的参数和权重在回复里。
原文

i am running stepfun's new step 3.7 flash on a dgx spark right now. 198b vision model, on a box that sits on a desk. here's how to save yourself about 3 hours of head scratching getting it loaded, because i already did the head scratching for you. the official README tells you you need stepfun's own llama.cpp fork. you don't, mainline ggml-org runs it fine, vision and all, at 64k. don't burn an hour building a fork just to find out the model loads without it. this is the one that'll actually eat your evening: the model is 104gb of your ~121gb unified pool, and the spark has no swap. ask for too much context and it doesn't crash clean, it silently thrashes, the kernel evicts model pages and re-reads them from disk in a loop while you stare at "loading" forever. the tell is read_bytes. if it climbs past the 104gb model size and keeps going while the process is pinned at 99% memory with no log progress, it's wedged. kill it, it won't recover on its own. with the vision projector loaded on a q8 kv cache, 64k is your ceiling on a 128gb box, push past it and it thrashes in the clip loader, the model plus KV plus vision buffers all fighting over the ~17gb you have left. want the full 256k context? drop vision and switch to a q4 kv cache, then it fits, 114 of 121gb. that's the trade their README doesn't spell out: the big context is real, it's just q4 cache, text-only. and it's a reasoning model. ask it something, if you get a blank reply, you didn't break it, you capped max_tokens too low and it spent the whole budget thinking. the answer's sitting in reasoning_content. give it room. that's the 3 hours. exact working flags and the weights in the reply.

在 X 查看被引用的帖子

来源:StepFun · x.com