Ant Ling· @AntLingAGI · X·· 2026-04-28精选AI 评分62
AI 导读
AntLingAGI与SGLang团队合作,正式推出Ling-2.6-flash(亦称Elephant-alpha)即时指令模型,并在SGLang平台上实现了首发支持。该模型总参数量达104B,但活跃参数仅7.4B,专为低延迟的智能体工作流优化,能够实现即时响应。它在编码、文档处理和智能体任务中展现出极高的token效率,所用token数量显著减少。尽管活跃参数较少,其模型质量仍与当前SOTA水平相当,兼具速度与执行力,适合需要快速响应的生产级智能体应用。团队强调,快速且稳定的推理是提升用户体验的关键。
推荐理由
104B 总参但只激活 7.4B,蚂蚁这步棋是冲着 Agent 场景的低延迟去的,做 Agent 产品的人值得跑一下看看实际体感。
正文 · 原文
🥳 It has always been our pleasure to work with the SGLang team, as we all believe in fast and stable inference is the key to our valuable users' experience.🫡
Hope you all enjoy Ling-2.6-flash (aka Elephant-alpha) 🐘⚡️⚡️
打满~ 打满~~ 😝
🎉 Meet Ling-2.6-flash from @AntLingAGI, an instant instruct model with 104B total params (7.4B active). Day-0 support is now live in SGLang! 1️⃣ Instant responses: tuned for low-latency agent workflows 2️⃣ SOTA-comparable quality with way fewer active params 3️⃣ High token efficiency: significantly fewer tokens used across coding, doc processing & agent tasks 4️⃣ Production-ready for agents that need speed + strong execution Run it now with SGLang!在 X 查看被引用的帖子
来源:Ant Ling · x.com