腾讯混元开源端到端投机解码框架 AngelSpec,支持训练与部署。在 Hy3-A21B 模型上,其 DFly 方案相比自回归解码实现 1.98–2.40 倍端到端加速,吞吐量比 DFlash 高 10.5–11.8%。训练代码及 Hy3-A21B MTP/DFly 草稿模型权重已开源。
做推理加速的可以试试这个端到端推测解码框架,腾讯开源了完整代码和权重,在 Hy3 上提速近 2.4 倍且吞吐更高,拿来即用。
🚀 We’ve open-sourced AngelSpec, an end-to-end speculative decoding framework supporting both training and deployment.
On Hy3-A21B, DFly delivers a 1.98–2.40× end-to-end speedup over autoregressive decoding across tested concurrency levels from 4 to 64, with 10.5–11.8% higher throughput than DFlash.
Training code and Hy3-A21B MTP/DFly drafter weights are now available:
GitHub: https://github.com/Tencent/AngelSpec
Paper: https://arxiv.org/abs/2607.25852
Docs: https://angelspec.readthedocs.io
Hugging Face: https://huggingface.co/collections/AngelSlim/angelspec
ModelScope: https://modelscope.cn/collections/AngelSlim/AngelSpec
#Hy3 #AngelSpec #OpenSource
来源:Tencent Hy · x.com