inclusionAI 发布 Ling-3.0-flash 的 DSpark 投机解码模型 Ling3-DSpark
inclusionAI/Ling-3.0-flash-dspark
inclusionAI 推出 Ling3-DSpark,一个为 Ling-3.0-flash 设计的 DSpark 投机解码模型,参数量 1.36B,通过置信度头动态选择草稿 token 数量。在 GSM8K 等九项基准上平均接受长度为 5.29,其中 GSM8K 达 6.40。模型经 SpecForge 训练,可通过 SGLang 部署。
该模型把投机解码的接受长度按工作负载分开报告,数学与代码任务明显高于对话类,部署时可据此对不同场景的加速收益有更实际预期。
Ling3-DSpark
A DSpark speculator for Ling3. DSpark extends DFlash with target-model auxiliary features and a confidence head that dynamically chooses the number of draft tokens. The model was trained with SpecForge and is served with SGLang.
Model specifications
- Target model: Ling-3.0-flash
- Draft parameters: 1,363,707,905 (1.36B)
- Draft weight dtype: BF16
- Hidden size: 2,560
- Transformer layers: 5 full-attention layers
- Attention: MHA with 32 query heads and 32 key/value heads
- Target auxiliary feature layers: 1, 11, 23, 29, 35
- Confidence head: vanilla Markov head, rank 256
- DSpark block size: 8 draft tokens (verify width 9, including the target bonus token)
- Maximum position embeddings: 262,144
Acceptance length
Acceptance length is the mean number of tokens accepted per speculative verification step, including the target bonus token.
| Workload | Acceptance length |
|---|---|
| GSM8K | 6.40 |
| MATH-500 | 6.29 |
| AIME 2025 | 5.56 |
| HumanEval | 6.57 |
| MBPP | 6.34 |
| LiveCodeBench | 5.33 |
| MT-Bench | 3.92 |
| Alpaca | 3.51 |
| Arena-Hard-v2 | 3.72 |
The macro mean across the nine workload means is 5.29.
Serving with SGLang
Use an SGLang version with DSPARK support. Replace the model paths and tensor-parallel size with values appropriate for your deployment:
sglang serve \
--trust-remote-code \
--model-path <LING3_MODEL_PATH> \
--tp-size <TP_SIZE> \
--speculative-algorithm DSPARK \
--speculative-draft-model-path <LING3_DSPARK_MODEL_PATH> \
......
1B params
来源:蚂蚁 inclusionAI:HuggingFace 新模型 · huggingface.co