团队推出 KPop,用于稳定大规模 MoE 模型的智能体强化学习训练。它用基于二元 KL 散度的自适应掩码机制,替代了此前 IcePop 方法中的固定比例掩码,能根据训练过程中的训练-推理不匹配程度动态调整。这一改进使得 Ring-2.6-1T 模型在无需修改基础设施或路由重放的情况下,仅通过纯 RL 训练,在 SWE-bench Verified 上取得了超过 76 分的成绩。
蚂蚁团队把 IcePop 升级成 KPop,从固定掩码变成自适应 KL 区域,思路很巧。Ring-2.6-1T 纯 RL 直接冲到 SWE-bench 76+,做 agentic RL 训练的同学值得翻一下博客。
从 IcePop 到 KPop——我们的团队持续推动大型 MoE 模型的 RL 训练稳定性。👇
KPop 用自适应二进制 KL 区域取代了固定比例掩码,该区域与每个 token 的固有噪声相匹配。更新更加稳健,实现稳定的长期智能体 RL。
Ring-2.6-1T → SWE-bench Verified 上 76+,纯 RL。
祝贺 @Jia__Guo 及团队!
好奇我们万亿级智能体基础模型背后的秘密武器吗?它来了!🥳 去年,我们发布了 IcePop,通过双边掩码来稳定 MoE RL。随着我们深入探索,意想不到的事情发生了:掩码比例下降了,而训练-推理不匹配却持续增长!😞 今年,我们推出 𝑲𝑷𝒐𝒑🪩,它用二元 KL 散度取代固定比例约束,自适应地掩码不合适的 token!掩码比例会随训练过程中训练-推理差距的波动而自适应调整,从而在长时程智能体 RL rollout 中保持策略优化稳定且有效。 通过这一简单改动,它使我们的 Ring-2.6-1T 在纯 RL 训练下于 SWE-bench-Verified 上取得超过 76 的成绩! 无需修改基础设施。无需 routing replay。只需一个参数,用 𝑲𝑷𝒐𝒑 为你的智能体 RL 赋能! 点击了解更多细节! 📜Blog: https://ringtech.notion.site/kpop
原文
Curious about the secret sauce behind our trillion-scale agentic foundation model? Here it comes!🥳 Last year, we released IcePop to stabilize MoE RL with double-sided masking. As we dive deeper, something unexpected happened: the masking ratio went down, while the training–inference mismatch continued to grow!😞 This year, we introduce 𝑲𝑷𝒐𝒑🪩, which replaces the fixed ratio constraint with the binary KL divergence to adaptively mask inappropriate tokens! The masking ratio adapts to fluctuations of the training–inference gap during training, keeping policy optimization stable and effective with long-horizon agentic RL rollouts. With this simple change, it enables our Ring-2.6-1T to achieve over 76 on the SWE-bench-Verified with pure RL training! No modifications to infrastructure. No routing replay. Just one parameter, power your agentic RL with 𝑲𝑷𝒐𝒑! Click to learn more about the details! 📜Blog: https://ringtech.notion.site/kpop
来源:Ant Ling · x.com