跳到正文
北京时间
原文
Hacker News 热门(buzzing.cc 中文翻译)· GreenGames·· 2026-04-21精选AI 评分70

我们在RTX 3090上运行Qwen3.5-27B,获得了207 tok/s的性能

AI 导读

开发者在单张RTX 3090显卡上成功运行Qwen3.5-27B模型,实现了每秒207个token的生成速度。该项目已在GitHub平台开源,展示了消费级硬件运行270亿参数大语言模型的高性能潜力。相关成果在Hacker News获得105个点赞,引发技术社区对本地大模型部署效率与优化方案的关注。

推荐理由

Lucebox 把 Qwen3.5-27B 在 3090 上推到了 207 tok/s,靠的是定制投机解码和 KV 压缩,做本地推理的开发者值得照着配一遍,虽然调试门槛不低。

正文 · 原文

Local LLM inference server built for speed. Custom kernels, speculative prefill & decoding. Each optimization in our engine is for specific model family and hardware target.

Inference Engine Optimizations

Each one is self-contained with setup instructions and benchmark notes.

Supported Models & Drafters

All speedups measured vs vendored llama.cpp (-fa 1, matching KV quant). Combined = geometric mean √(TTFT × decode) where both phases benched; otherwise the single-phase speedup. Drafters published on huggingface.co/Lucebox.

-fa 1

Model Speedup Qwen 3.5-0.8B (Megakernel) ~2× Qwen 3.5-27B + DDTree 3.43× Qwen 3.6-27B + PFlash ~5.6× Qwen 3.6-27B + DDTree 4.84× Laguna-XS.2 33B + PFlash 5.4× @128K Qwen 3.5-27B HIP ~2.6× Gemma-4-26B-A4B 1.31× Drafter Phase Qwen3.6-27B decode gemma-4-26B-A4B decode gemma-4-31B decode Qwen3-0.6B prefill

Model Speedup Qwen 3.5-0.8B (Megakernel) ~2× Qwen 3.5-27B + DDTree 3.43× Qwen 3.6-27B + PFlash ~5.6× Qwen 3.6-27B + DDTree 4.84× Laguna-XS.2 33B + PFlash 5.4× @128K Qwen 3.5-27B HIP ~2.6× Gemma-4-26B-A4B 1.31×

Drafter Phase Qwen3.6-27B decode gemma-4-26B-A4B decode gemma-4-31B decode Qwen3-0.6B prefill

Qwen3.6-27B

gemma-4-26B-A4B

gemma-4-31B

Qwen3-0.6B

Tested Machines (GPU/APU)

Reference target: RTX 3090 (Ampere sm_86) — all headline numbers. Other NVIDIA archs auto-detected by CMake / setup.py; AMD HIP backend separate (Strix Halo section).

setup.py

Arch GPU Min CUDA / ROCm Status Bench Ampere sm_86 RTX 3090, A-series CUDA 12.0 ✅ reference megakernel · dflash Blackwell sm_120 RTX 5090 CUDA 12.8 ✅ 205 tok/s, 4.84× ↗ Blackwell sm_121 DGX Spark / GB10 CUDA 12.9 ✅ megakernel NVFP4 ↗ Turing sm_75 RTX 2080 Ti CUDA 12.0 ✅ 53 tok/s DFlash ↗ Ada sm_89 RTX 40xx CUDA 12.0 🟡 community WSL2 bench ↗ — Blackwell sm_110 Jetson AGX Thor CUDA 13.0 🟡 builds, unbenched — Volta sm_70 / Pascal sm_61 V100, P40 CUDA 12.0 🟡 fallback paths, unbenched — RDNA3.5 gfx1151 Ryzen AI MAX+ 395 / Strix Halo ROCm 6+ ✅ 37 tok/s HIP ↗ RDNA3 gfx1100 Radeon RX 7900 XTX ROCm 6+ ✅ 50 tok/s HIP ↗

sm_86

sm_120

sm_121

sm_75

sm_89

sm_110

sm_70

sm_61

gfx1151

gfx1100

server/ (DFlash) builds with CMake 3.18+ and --recurse-submodules for Luce-Org/llama.cpp@luce-dflash — no PyTorch needed. optimizations/megakernel/ is the only component requiring PyTorch 2.0+ (CUDAExtension links against torch C++ libs). Power-tune: sudo nvidia-smi -pl 220 (3090 sweet spot, re-sweep for other cards).

server/

--recurse-submodules

Luce-Org/llama.cpp@luce-dflash

optimizations/megakernel/

sudo nvidia-smi -pl 220

Quick Start On Harnesses

harness/ contains RTX 3090 client launchers and regression tests for Lucebox server compatibility. Run Lucebox inside Claude Code, Codex, OpenCode, Hermes, Pi, OpenClaw, or Open WebUI, or check if a server change still works with those clients.

harness/

Client Launcher Claude Code run_claude_code.sh Codex run_codex.sh OpenCode run_opencode.sh Hermes run_hermes.sh Pi run_pi.sh OpenClaw run_openclaw.sh Open WebUI run_openwebui.sh

Client Launcher Claude Code run_claude_code.sh Codex run_codex.sh OpenCode run_opencode.sh Hermes run_hermes.sh Pi run_pi.sh OpenClaw run_openclaw.sh Open WebUI run_openwebui.sh

run_claude_code.sh

run_codex.sh

run_opencode.sh

run_hermes.sh

run_pi.sh

run_openclaw.sh

run_openwebui.sh

All launchers spawn the native C++ HTTP server (dflash_server). Override defaults via env vars:

dflash_server

DFLASH_SERVER_BIN=server/build/dflash_server \ DFLASH_TARGET=server/models/Qwen3.6-27B-Q4_K_M.gguf \ DFLASH_DRAFT=server/models/draft/dflash-draft-3.6-q4_k_m.gguf \ MAX_CTX=32768 BUDGET=22 VERIFY_MODE=ddtree \ harness/clients/run_codex.sh

For no-draft targets such as Gemma, set only DFLASH_TARGET or pass DRAFT=none; the harness will not attach the default Qwen draft to a custom target.

DFLASH_TARGET

DRAFT=none

Launcher scripts install missing real-client CLIs automatically under .harness-work/. To preinstall them yourself:

.harness-work/

python3 harness/client_test_runner.py install --clients codex,hermes,openwebui

For direct TPS/TTFT numbers against a running server:

python3 harness/client_test_runner.py bench \ --url http://127.0.0.1:8000 \ --suite he,agent \ --n-sample 3

Quick Start With Docker

Prebuilt images on GHCR track main. No CUDA toolkit or build needed. Pull the image, mount weights and serve. OpenAI-compatible API on :8000.

main

GPU Image tag NVIDIA (CUDA 12+) :cuda12 AMD (ROCm 6+) :rocm Drop a GGUF model target into server/models/ first, then :8000/v1/chat/completions. Full tutorial in the Docker blog.

GPU Image tag NVIDIA (CUDA 12+) :cuda12 AMD (ROCm 6+) :rocm

:cuda12

:rocm

Drop a GGUF model target into server/models/ first, then :8000/v1/chat/completions. Full tutorial in the Docker blog.

server/models/

:8000/v1/chat/completions

Install and run:

1. Pull the image for your GPU docker pull ghcr.io/luce-org/lucebox-hub:cuda12 # NVIDIA docker pull ghcr.io/luce-org/lucebox-hub:rocm # AMD # 2. Download a target model into server/models/ and the DFlash draft # into server/models/draft/ (the entrypoint only auto-discovers the # draft there; without it the server runs slower, target-only) hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf \ --local-dir server/models/ hf download Lucebox/Qwen3.6-27B-DFlash-GGUF dflash-draft-3.6-q4_k_m.gguf \ --local-dir server/models/draft/ # 3a. NVIDIA (CUDA 12+) docker run --rm --gpus all -p 8000:8080 \ -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:cuda12 # 3b. AMD (ROCm 6+, Strix Halo / RX 7900) docker run --rm --device /dev/kfd --device /dev/dri \ --group-add video --group-add render --security-opt seccomp=unconfined \ -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:rocm

Then hit :8000/v1/chat/completions (OpenAI-compatible).

:8000/v1/chat/completions

Run the Server

Default: Qwen 3.6-27B Q4_K_M target + Lucebox Q4_K_M DFlash drafter on RTX 3090. DDTree budget=22, TQ3_0 KV cache, sliding FA window 2048. OpenAI-compatible HTTP on :8000.

build (CUDA 12+, CMake 3.18+) git clone --recurse-submodules https://github.com/Luce-Org/lucebox-hub && cd lucebox-hub cmake -B server/build -S server -DCMAKE_BUILD_TYPE=Release cmake --build server/build --target dflash_server -j # default weights (~18 GB) hf download unsloth/Qwen3.6-27B-GGUF Qwen3.6-27B-Q4_K_M.gguf --local-dir server/models/ hf download Lucebox/Qwen3.6-27B-DFlash-GGUF dflash-draft-3.6-q4_k_m.gguf --local-dir server/models/draft/ # run (TQ3_0 KV auto-enabled; set =0 to disable) DFLASH27B_KV_TQ3=1 \ ./server/build/dflash_server server/models/Qwen3.6-27B-Q4_K_M.gguf \ --draft server/models/draft/dflash-draft-3.6-q4_k_m.gguf \ --ddtree --ddtree-budget 22 --fa-window 2048 --port 8000

Server flags

Core

Flag Default Effect --draft — DFlash draft GGUF, required for speculative decode --port N 8000 HTTP port --host H 127.0.0.1 Bind address --max-ctx N auto-fit KV cache size; oversizing slows prefill (FA stride over unused KV) --max-tokens N model-card Generation cap --model-name S filename OpenAI model field --chat-template-file autodetect Override Jinja template

--draft

--port N

--host H

127.0.0.1

--max-ctx N

--max-tokens N

--model-name S

model

--chat-template-file

Decode (DFlash + DDTree)

Flag Default Effect --ddtree off (chain) Enable tree verify --ddtree-budget N 22 Tree size. 22 on 3090 (default), 40 on 5090, re-sweep on GB10 --fa-window N 2048 Sliding FA window; 0 = full attention --draft-residency {auto,persistent,request-scoped} auto When draft weights are evicted from VRAM. request-scoped parks/frees them after each request's draft work (frees VRAM for the target on tight GPUs); persistent keeps them resident across requests; auto preserves current behavior while honoring the low-VRAM / --lazy-draft hint. Reported at /props.runtime.draft_residency. --lazy-draft off Legacy alias for --draft-residency=request-scoped (defer draft load until first request, release after)

--ddtree

--ddtree-budget N

--fa-window N

--draft-residency {auto,persistent,request-scoped}

auto

request-scoped

persistent

auto

--lazy-draft

/props.runtime.draft_residency

--lazy-draft

--draft-residency=request-scoped

Prefill compression (PFlash)

Flag / env Default Effect --prefill-compression {off,auto,always} off When to score+compress the prompt --prefill-threshold N 32000 In auto, the prompt-token count above which a single-shot prompt is compressed. Also the per-message minimum that an aged message must exceed before FlowKV compresses it on multi-turn requests. Lower it (e.g. 1024) if you want FlowKV to act on shorter history. --prefill-keep-ratio F 0.05 Fraction of source tokens kept (0.02 @128K, 0.10 @32K) --prefill-curve T:R [T:R ...] off (flat keep-ratio) Piecewise keep-ratio curve, linear-interpolated over (tokens, ratio) breakpoints, e.g. 10000:0.5 40000:0.2 100000:0.1 (2× compression @10K, 5× @40K, 10× @100K+). Overrides --prefill-keep-ratio; per-session bandit override still wins. --prefill-drafter required if on Drafter weights (Qwen3-0.6B BF16 GGUF) --prefill-skip-park off Keep drafter resident across requests (more VRAM, faster) PFLASH_FREEZE_HOT_WINDOW=N 2 FlowKV: how many of the most recent messages stay verbatim. Everything older than this window (but after the system prompt) is compressed once and cached. Larger = more recent context kept uncompressed. DFLASH_FP_USE_BSA=1 0 Dispatch sparse FA through BSA (sm_80+); required for headline 10.4× DFLASH_FP_ALPHA=0.85 0.12 Block-selection threshold; higher = stricter = fewer K-blocks DFLASH_FP_PROFILE=1 0 Per-stage timing log

--prefill-compression {off,auto,always}

off

--prefill-threshold N

auto

--prefill-keep-ratio F

0.05

--prefill-curve T:R [T:R ...]

(tokens, ratio)

10000:0.5 40000:0.2 100000:0.1

--prefill-keep-ratio

--prefill-drafter

--prefill-skip-park

PFLASH_FREEZE_HOT_WINDOW=N

DFLASH_FP_USE_BSA=1

DFLASH_FP_ALPHA=0.85

0.12

DFLASH_FP_PROFILE=1

When compression is on, the request path picks one of three modes automatically, so they never stack: the first turn is sent verbatim (the system prompt stays as a stable cache anchor), multi-turn continuations use FlowKV (only the aged history is compressed, recent turns kept verbatim, so the disk prefix cache from --prefix-cache-slots keeps hitting), and a single oversized prompt with no prior turns uses whole-prompt PFlash. With --prefill-compression off the request path is identical to a build without compression.

--prefix-cache-slots

--prefill-compression off

KV cache

Flag / env Default Effect --cache-type-k / --cache-type-v env-driven Per-side quant override: f16,bf16,q4_0,q4_1,q5_0,q5_1,q8_0,tq3_0 DFLASH27B_KV_TQ3=1 (default) Preset TQ3_0 K+V (3.5 bpv, fits 256K @ 24 GB) DFLASH27B_KV_Q4=1 off Q4_0 K+V (4.5 bpv, legacy, ~128K ceiling) --prefix-cache-slots N — Live prefix-cache slot count --kv-cache-dir — Persist prefix cache to disk --kv-cache-budget N — On-disk cache size cap

--cache-type-k

--cache-type-v

f16,bf16,q4_0,q4_1,q5_0,q5_1,q8_0,tq3_0

DFLASH27B_KV_TQ3=1

DFLASH27B_KV_Q4=1

--prefix-cache-slots N

--kv-cache-dir

--kv-cache-budget N

Bounded KV residency (KVFlash)

Pages the attention KV cache through a fixed pool of GPU slots; cold 64-token chunks live in host RAM, bit-exact and recallable. Decode speed stops depending on context length and resident KV stays pool-sized at any context. Off by default; works on every model family. Drafter-scored residency is the default on every family: the server finds the Qwen3-0.6B drafter next to the model (or via --prefill-drafter) and lazy-loads it as the relevance scorer that decides which chunks stay resident — non-qwen targets (laguna, gemma4) bridge the tokenizer gap by re-tokenizing the context text for the drafter. LRU is the fallback when no drafter is present, or the explicit choice via --kvflash-policy lru. Per-model numbers in Luce KVFlash →.

--prefill-drafter

--kvflash-policy lru

Flag / env Default Effect --kvflash <tokens|auto> off Resident pool size. auto sizes from the GPU: half of free VRAM after weights and reserves, at the model's KV density, capped where decode speed stays near the flat optimum (default 16384, override DFLASH_KVFLASH_MAX_POOL) and at --max-ctx. Explicit values are rounded to 256, clamped to --max-ctx, floored at the protected minimum so eviction always has a victim. --kvflash-policy {drafter,lru} drafter Residency policy. lru opts out of the drafter probe/load (recency-only paging, no extra VRAM). --kvflash-tau N 64 Reselect interval floor (drafter policy only); the effective interval grows with history to cap rescore overhead. DFLASH_KVFLASH=N off Env equivalent of --kvflash. DFLASH_KVFLASH_TAU=N 64 Env equivalent of --kvflash-tau.

--kvflash <tokens|auto>

auto

DFLASH_KVFLASH_MAX_POOL

--max-ctx

--max-ctx

--kvflash-policy {drafter,lru}

drafter

lru

--kvflash-tau N

DFLASH_KVFLASH=N

--kvflash

DFLASH_KVFLASH_TAU=N

--kvflash-tau

Thinking budget

Flag Default Effect --think-max-tokens N model-card Max tokens inside … --default-max-tokens N model-card Default response cap --hard-limit-reply-budget N 4096 Hard ceiling; injects close near limit --reasoning-effort-{low,medium,high,x-high,max} N model-card OpenAI-style effort tiers

--think-max-tokens N

--default-max-tokens N

--hard-limit-reply-budget N

--reasoning-effort-{low,medium,high,x-high,max} N

Multi-GPU / IPC

Flag / env Default Effect --target-device cuda:0 Target backend (e.g. cuda:0, hip:0) --draft-device same as target Draft backend; mixed backend needs --draft-ipc-bin --target-gpu N 0 Target GPU index --draft-gpu N same as target Draft GPU index; offload draft to a second GPU --target-devices / --target-layer-split single GPU Layer-split target across GPUs --draft-ipc-bin — Out-of-process draft binary (mixed CUDA/HIP) --peer-access off Enable P2P between target GPUs --chunk N backend default Prefill ubatch size --no-cors CORS on Disable CORS headers DFLASH_TARGET_GPU=N 0 Env var equivalent of --target-gpu DFLASH_DRAFT_GPU=N same as target Env var equivalent of --draft-gpu

--target-device

cuda:0

cuda:0

hip:0

--draft-device

--draft-ipc-bin

--target-gpu N

--draft-gpu N

--target-devices

--target-layer-split

--draft-ipc-bin

--peer-access

--chunk N

--no-cors

DFLASH_TARGET_GPU=N

--target-gpu

DFLASH_DRAFT_GPU=N

来源:Hacker News 热门(buzzing.cc 中文翻译) · github.com