我们在RTX 3090上运行Qwen3.5-27B,获得了207 tok/s的性能
开发者在单张RTX 3090显卡上成功运行Qwen3.5-27B模型,实现了每秒207个token的生成速度。该项目已在GitHub平台开源,展示了消费级硬件运行270亿参数大语言模型的高性能潜力。相关成果在Hacker News获得105个点赞,引发技术社区对本地大模型部署效率与优化方案的关注。
Lucebox 把 Qwen3.5-27B 在 3090 上推到了 207 tok/s,靠的是定制投机解码和 KV 压缩,做本地推理的开发者值得照着配一遍,虽然调试门槛不低。
Local LLM inference server built for speed. Custom kernels, speculative prefill & decoding. Each optimization in our engine is for specific model family and hardware target.
Inference Engine Optimizations
Each one is self-contained with setup instructions and benchmark notes.
Supported Models & Drafters
All speedups measured vs vendored llama.cpp (-fa 1, matching KV quant). Combined = geometric mean √(TTFT × decode) where both phases benched; otherwise the single-phase speedup. Drafters published on huggingface.co/Lucebox.
-fa 1
Model Speedup Qwen 3.5-0.8B (Megakernel) ~2× Qwen 3.5-27B + DDTree 3.43× Qwen 3.6-27B + PFlash ~5.6× Qwen 3.6-27B + DDTree 4.84× Laguna-XS.2 33B + PFlash 5.4× @128K Qwen 3.5-27B HIP ~2.6× Gemma-4-26B-A4B 1.31× Drafter Phase Qwen3.6-27B decode gemma-4-26B-A4B decode gemma-4-31B decode Qwen3-0.6B prefill
Model Speedup Qwen 3.5-0.8B (Megakernel) ~2× Qwen 3.5-27B + DDTree 3.43× Qwen 3.6-27B + PFlash ~5.6× Qwen 3.6-27B + DDTree 4.84× Laguna-XS.2 33B + PFlash 5.4× @128K Qwen 3.5-27B HIP ~2.6× Gemma-4-26B-A4B 1.31×
Drafter Phase Qwen3.6-27B decode gemma-4-26B-A4B decode gemma-4-31B decode Qwen3-0.6B prefill
Qwen3.6-27B
gemma-4-26B-A4B
gemma-4-31B
Qwen3-0.6B
Tested Machines (GPU/APU)
Reference target: RTX 3090 (Ampere sm_86) — all headline numbers. Other NVIDIA archs auto-detected by CMake / setup.py; AMD HIP backend separate (Strix Halo section).
setup.py
Arch GPU Min CUDA / ROCm Status Bench Ampere sm_86 RTX 3090, A-series CUDA 12.0 ✅ reference megakernel · dflash Blackwell sm_120 RTX 5090 CUDA 12.8 ✅ 205 tok/s, 4.84× ↗ Blackwell sm_121 DGX Spark / GB10 CUDA 12.9 ✅ megakernel NVFP4 ↗ Turing sm_75 RTX 2080 Ti CUDA 12.0 ✅ 53 tok/s DFlash ↗ Ada sm_89 RTX 40xx CUDA 12.0 🟡 community WSL2 bench ↗ — Blackwell sm_110 Jetson AGX Thor CUDA 13.0 🟡 builds, unbenched — Volta sm_70 / Pascal sm_61 V100, P40 CUDA 12.0 🟡 fallback paths, unbenched — RDNA3.5 gfx1151 Ryzen AI MAX+ 395 / Strix Halo ROCm 6+ ✅ 37 tok/s HIP ↗ RDNA3 gfx1100 Radeon RX 7900 XTX ROCm 6+ ✅ 50 tok/s HIP ↗
sm_86
sm_120
sm_121
sm_75
sm_89
sm_110
sm_70
sm_61
gfx1151
gfx1100
server/ (DFlash) builds with CMake 3.18+ and --recurse-submodules for Luce-Org/llama.cpp@luce-dflash — no PyTorch needed. optimizations/megakernel/ is the only component requiring PyTorch 2.0+ (CUDAExtension links against torch C++ libs). Power-tune: sudo nvidia-smi -pl 220 (3090 sweet spot, re-sweep for other cards).
server/
--recurse-submodules
Luce-Org/llama.cpp@luce-dflash
optimizations/megakernel/
sudo nvidia-smi -pl 220
Quick Start On Harnesses
harness/ contains RTX 3090 client launchers and regression tests for Lucebox server compatibility. Run Lucebox inside Claude Code, Codex, OpenCode, Hermes, Pi, OpenClaw, or Open WebUI, or check if a server change still works with those clients.
harness/
Client Launcher Claude Code run_claude_code.sh Codex run_codex.sh OpenCode run_opencode.sh Hermes run_hermes.sh Pi run_pi.sh OpenClaw run_openclaw.sh Open WebUI run_openwebui.sh
Client Launcher Claude Code run_claude_code.sh Codex run_codex.sh OpenCode run_opencode.sh Hermes run_hermes.sh Pi run_pi.sh OpenClaw run_openclaw.sh Open WebUI run_openwebui.sh
run_claude_code.sh
run_codex.sh
run_opencode.sh
run_hermes.sh
run_pi.sh
run_openclaw.sh
run_openwebui.sh
All launchers spawn the native C++ HTTP server (dflash_server). Override defaults via env vars:
dflash_server
DFLASH_SERVER_BIN=server/build/dflash_server \ DFLASH_TARGET=server/models/Qwen3.6-27B-Q4_K_M.gguf \ DFLASH_DRAFT=server/models/draft/dflash-draft-3.6-q4_k_m.gguf \ MAX_CTX=32768 BUDGET=22 VERIFY_MODE=ddtree \ harness/clients/run_codex.sh
For no-draft targets such as Gemma, set only DFLASH_TARGET or pass DRAFT=none; the harness will not attach the default Qwen draft to a custom target.
DFLASH_TARGET
DRAFT=none
Launcher scripts install missing real-client CLIs automatically under .harness-work/. To preinstall them yourself:
.harness-work/
python3 harness/client_test_runner.py install --clients codex,hermes,openwebui
For direct TPS/TTFT numbers against a running server:
python3 harness/client_test_runner.py bench \ --url http://127.0.0.1:8000 \ --suite he,agent \ --n-sample 3
Quick Start With Docker
Prebuilt images on GHCR track main. No CUDA toolkit or build needed. Pull the image, mount weights and serve. OpenAI-compatible API on :8000.
main
GPU Image tag NVIDIA (CUDA 12+) :cuda12 AMD (ROCm 6+) :rocm Drop a GGUF model target into server/models/ first, then :8000/v1/chat/completions. Full tutorial in the Docker blog.
GPU Image tag NVIDIA (CUDA 12+) :cuda12 AMD (ROCm 6+) :rocm
:cuda12
:rocm
Drop a GGUF model target into server/models/ first, then :8000/v1/chat/completions. Full tutorial in the Docker blog.
server/models/
:8000/v1/chat/completions
Install and run:
1. Pull the image for your GPU docker pull ghcr.io/luce-org/lucebox-hub:cuda12 # NVIDIA docker pull ghcr.io/luce-org/lucebox-hub:rocm # AMD # 2. Download a target model into server/models/ and the DFlash draft # into server/models/draft/ (the entrypoint only auto-discovers the # draft there; without it the server runs slower, target-only) hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf \ --local-dir server/models/ hf download Lucebox/Qwen3.6-27B-DFlash-GGUF dflash-draft-3.6-q4_k_m.gguf \ --local-dir server/models/draft/ # 3a. NVIDIA (CUDA 12+) docker run --rm --gpus all -p 8000:8080 \ -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:cuda12 # 3b. AMD (ROCm 6+, Strix Halo / RX 7900) docker run --rm --device /dev/kfd --device /dev/dri \ --group-add video --group-add render --security-opt seccomp=unconfined \ -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:rocm
Then hit :8000/v1/chat/completions (OpenAI-compatible).
:8000/v1/chat/completions
Run the Server
Default: Qwen 3.6-27B Q4_K_M target + Lucebox Q4_K_M DFlash drafter on RTX 3090. DDTree budget=22, TQ3_0 KV cache, sliding FA window 2048. OpenAI-compatible HTTP on :8000.
build (CUDA 12+, CMake 3.18+) git clone --recurse-submodules https://github.com/Luce-Org/lucebox-hub && cd lucebox-hub cmake -B server/build -S server -DCMAKE_BUILD_TYPE=Release cmake --build server/build --target dflash_server -j # default weights (~18 GB) hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir server/models/ hf download Lucebox/Qwen3.6-27B-DFlash-GGUF dflash-draft-3.6-q4_k_m.gguf --local-dir server/models/draft/ # run (TQ3_0 KV auto-enabled; set =0 to disable) DFLASH27B_KV_TQ3=1 \ ./server/build/dflash_server server/models/Qwen3.6-27B-Q4_K_M.gguf \ --draft server/models/draft/dflash-draft-3.6-q4_k_m.gguf \ --ddtree --ddtree-budget 22 --fa-window 2048 --port 8000
Server flags
Core
Flag Default Effect --draft — DFlash draft GGUF, required for speculative decode --port N 8000 HTTP port --host H 127.0.0.1 Bind address --max-ctx N auto-fit KV cache size; oversizing slows prefill (FA stride over unused KV) --max-tokens N model-card Generation cap --model-name S filename OpenAI model field --chat-template-file autodetect Override Jinja template
--draft
--port N
--host H
127.0.0.1
--max-ctx N
--max-tokens N
--model-name S
model
--chat-template-file
Decode (DFlash + DDTree)
Flag Default Effect --ddtree off (chain) Enable tree verify --ddtree-budget N 22 Tree size. 22 on 3090 (default), 40 on 5090, re-sweep on GB10 --fa-window N 2048 Sliding FA window; 0 = full attention --draft-residency {auto,persistent,request-scoped} auto When draft weights are evicted from VRAM. request-scoped parks/frees them after each request's draft work (frees VRAM for the target on tight GPUs); persistent keeps them resident across requests; auto preserves current behavior while honoring the low-VRAM / --lazy-draft hint. Reported at /props.runtime.draft_residency. --lazy-draft off Legacy alias for --draft-residency=request-scoped (defer draft load until first request, release after)
--ddtree
--ddtree-budget N
--fa-window N
--draft-residency {auto,persistent,request-scoped}
auto
request-scoped
persistent
auto
--lazy-draft
/props.runtime.draft_residency
--lazy-draft
--draft-residency=request-scoped
Prefill compression (PFlash)
Flag / env Default Effect --prefill-compression {off,auto,always} off When to score+compress the prompt --prefill-threshold N 32000 In auto, the prompt-token count above which a single-shot prompt is compressed. Also the per-message minimum that an aged message must exceed before FlowKV compresses it on multi-turn requests. Lower it (e.g. 1024) if you want FlowKV to act on shorter history. --prefill-keep-ratio F 0.05 Fraction of source tokens kept (0.02 @128K, 0.10 @32K) --prefill-curve T:R [T:R ...] off (flat keep-ratio) Piecewise keep-ratio curve, linear-interpolated over (tokens, ratio) breakpoints, e.g. 10000:0.5 40000:0.2 100000:0.1 (2× compression @10K, 5× @40K, 10× @100K+). Overrides --prefill-keep-ratio; per-session bandit override still wins. --prefill-drafter required if on Drafter weights (Qwen3-0.6B BF16 GGUF) --prefill-skip-park off Keep drafter resident across requests (more VRAM, faster) PFLASH_FREEZE_HOT_WINDOW=N 2 FlowKV: how many of the most recent messages stay verbatim. Everything older than this window (but after the system prompt) is compressed once and cached. Larger = more recent context kept uncompressed. DFLASH_FP_USE_BSA=1 0 Dispatch sparse FA through BSA (sm_80+); required for headline 10.4× DFLASH_FP_ALPHA=0.85 0.12 Block-selection threshold; higher = stricter = fewer K-blocks DFLASH_FP_PROFILE=1 0 Per-stage timing log
--prefill-compression {off,auto,always}
off
--prefill-threshold N
auto
--prefill-keep-ratio F
0.05
--prefill-curve T:R [T:R ...]
(tokens, ratio)
10000:0.5 40000:0.2 100000:0.1
--prefill-keep-ratio
--prefill-drafter
--prefill-skip-park
PFLASH_FREEZE_HOT_WINDOW=N
DFLASH_FP_USE_BSA=1
DFLASH_FP_ALPHA=0.85
0.12
DFLASH_FP_PROFILE=1
When compression is on, the request path picks one of three modes automatically, so they never stack: the first turn is sent verbatim (the system prompt stays as a stable cache anchor), multi-turn continuations use FlowKV (only the aged history is compressed, recent turns kept verbatim, so the disk prefix cache from --prefix-cache-slots keeps hitting), and a single oversized prompt with no prior turns uses whole-prompt PFlash. With --prefill-compression off the request path is identical to a build without compression.
--prefix-cache-slots
--prefill-compression off
KV cache
Flag / env Default Effect --cache-type-k / --cache-type-v env-driven Per-side quant override: f16,bf16,q4_0,q4_1,q5_0,q5_1,q8_0,tq3_0 DFLASH27B_KV_TQ3=1 (default) Preset TQ3_0 K+V (3.5 bpv, fits 256K @ 24 GB) DFLASH27B_KV_Q4=1 off Q4_0 K+V (4.5 bpv, legacy, ~128K ceiling) --prefix-cache-slots N — Live prefix-cache slot count --kv-cache-dir — Persist prefix cache to disk --kv-cache-budget N — On-disk cache size cap
--cache-type-k
--cache-type-v
f16,bf16,q4_0,q4_1,q5_0,q5_1,q8_0,tq3_0
DFLASH27B_KV_TQ3=1
DFLASH27B_KV_Q4=1
--prefix-cache-slots N
--kv-cache-dir
--kv-cache-budget N
Bounded KV residency (KVFlash)
Pages the attention KV cache through a fixed pool of GPU slots; cold 64-token chunks live in host RAM, bit-exact and recallable. Decode speed stops depending on context length and resident KV stays pool-sized at any context. Off by default; works on every model family. Drafter-scored residency is the default on every family: the server finds the Qwen3-0.6B drafter next to the model (or via --prefill-drafter) and lazy-loads it as the relevance scorer that decides which chunks stay resident — non-qwen targets (laguna, gemma4) bridge the tokenizer gap by re-tokenizing the context text for the drafter. LRU is the fallback when no drafter is present, or the explicit choice via --kvflash-policy lru. Per-model numbers in Luce KVFlash →.
--prefill-drafter
--kvflash-policy lru
Flag / env Default Effect --kvflash <tokens|auto> off Resident pool size. auto sizes from the GPU: half of free VRAM after weights and reserves, at the model's KV density, capped where decode speed stays near the flat optimum (default 16384, override DFLASH_KVFLASH_MAX_POOL) and at --max-ctx. Explicit values are rounded to 256, clamped to --max-ctx, floored at the protected minimum so eviction always has a victim. --kvflash-policy {drafter,lru} drafter Residency policy. lru opts out of the drafter probe/load (recency-only paging, no extra VRAM). --kvflash-tau N 64 Reselect interval floor (drafter policy only); the effective interval grows with history to cap rescore overhead. DFLASH_KVFLASH=N off Env equivalent of --kvflash. DFLASH_KVFLASH_TAU=N 64 Env equivalent of --kvflash-tau.
--kvflash <tokens|auto>
auto
DFLASH_KVFLASH_MAX_POOL
--max-ctx
--max-ctx
--kvflash-policy {drafter,lru}
drafter
lru
--kvflash-tau N
DFLASH_KVFLASH=N
--kvflash
DFLASH_KVFLASH_TAU=N
--kvflash-tau
Thinking budget
Flag Default Effect --think-max-tokens N model-card Max tokens inside … --default-max-tokens N model-card Default response cap --hard-limit-reply-budget N 4096 Hard ceiling; injects close near limit --reasoning-effort-{low,medium,high,x-high,max} N model-card OpenAI-style effort tiers
--think-max-tokens N
--default-max-tokens N
--hard-limit-reply-budget N
--reasoning-effort-{low,medium,high,x-high,max} N
Multi-GPU / IPC
Flag / env Default Effect --target-device cuda:0 Target backend (e.g. cuda:0, hip:0) --draft-device same as target Draft backend; mixed backend needs --draft-ipc-bin --target-gpu N 0 Target GPU index --draft-gpu N same as target Draft GPU index; offload draft to a second GPU --target-devices / --target-layer-split single GPU Layer-split target across GPUs --draft-ipc-bin — Out-of-process draft binary (mixed CUDA/HIP) --peer-access off Enable P2P between target GPUs --chunk N backend default Prefill ubatch size --no-cors CORS on Disable CORS headers DFLASH_TARGET_GPU=N 0 Env var equivalent of --target-gpu DFLASH_DRAFT_GPU=N same as target Env var equivalent of --draft-gpu
--target-device
cuda:0
cuda:0
hip:0
--draft-device
--draft-ipc-bin
--target-gpu N
--draft-gpu N
--target-devices
--target-layer-split
--draft-ipc-bin
--peer-access
--chunk N
--no-cors
DFLASH_TARGET_GPU=N
--target-gpu
DFLASH_DRAFT_GPU=N
来源:Hacker News 热门(buzzing.cc 中文翻译) · github.com