跳到正文
北京时间
原文
StepFun· @StepFun_ai · X·· 2026-05-31精选AI 评分80
AI 导读

阶跃星辰发布了Step 3.7 Flash,这是一款198B参数的视觉模型,旨在DGX Spark等桌面设备上运行。用户实测表明,128GB统一内存是运行门槛,模型占用约104GB。部署无需官方专用llama.cpp分支,主线版本即可。在上下文长度上存在权衡:启用视觉功能时,基于q8 KV cache的64K为上限;若要使用最高256K上下文,则需禁用视觉并切换至q4 KV cache,此时模型与缓存共占约114GB内存。该模型是推理模型,思考过程可能消耗大量max_tokens,需注意设置。

推荐理由

把 198B 的视觉模型塞进一台桌面盒子,还跑通了,这本身就是个小里程碑。更关键的是,这篇实战直接帮你绕开了三个大坑,省下的三小时够你喝杯咖啡慢慢试了。

正文 · 原文

A 198B vision model, running on a box that sits on a desk. This is what we built Step 3.7 Flash for.

Brilliant breakdown @sudoingX — saved everyone a few hours of head-scratching 🎉

引用Sudo su@sudoingX
i am running stepfun's new step 3.7 flash on a dgx spark right now. 198b vision model, on a box that sits on a desk. here's how to save yourself about 3 hours of head scratching getting it loaded, because i already did the head scratching for you. the official README tells you you need stepfun's own llama.cpp fork. you don't, mainline ggml-org runs it fine, vision and all, at 64k. don't burn an hour building a fork just to find out the model loads without it. this is the one that'll actually eat your evening: the model is 104gb of your ~121gb unified pool, and the spark has no swap. ask for too much context and it doesn't crash clean, it silently thrashes, the kernel evicts model pages and re-reads them from disk in a loop while you stare at "loading" forever. the tell is read_bytes. if it climbs past the 104gb model size and keeps going while the process is pinned at 99% memory with no log progress, it's wedged. kill it, it won't recover on its own. with the vision projector loaded on a q8 kv cache, 64k is your ceiling on a 128gb box, push past it and it thrashes in the clip loader, the model plus KV plus vision buffers all fighting over the ~17gb you have left. want the full 256k context? drop vision and switch to a q4 kv cache, then it fits, 114 of 121gb. that's the trade their README doesn't spell out: the big context is real, it's just q4 cache, text-only. and it's a reasoning model. ask it something, if you get a blank reply, you didn't break it, you capped max_tokens too low and it spent the whole budget thinking. the answer's sitting in reasoning_content. give it room. that's the 3 hours. exact working flags and the weights in the reply.
在 X 查看被引用的帖子

来源:StepFun · x.com