跳到正文
北京时间
原文
MiniMax (official)· @MiniMax_AI · X·· 2026-06-13精选AI 评分82
AI 导读

MiniMax M3 发布,具备前沿编码与智能体能力,原生图像视频输入和计算机使用,1M-token 上下文。核心采用 MSA 稀疏注意力:每个 query 评分 128-token KV 块,仅对 top 块做注意力。vLLM 当日即支持 M3,包括专用 MSA prefill/decode 核、前缀缓存与分块 prefill、BF16 和 MXFP8 检查点、Hopper 与 Blackwell 的 MoE 后端,并在 NVIDIA 与 AMD 硬件上验证。同时支持原生多模态输入、工具调用、推理解析和思考模式控制等智能体工作负载。

推荐理由

M3把1M上下文从‘理论上能做’变成了‘今天就能部署’,MSA稀疏注意力是关键,开源社区和推理框架的深度合作值得关注。

正文 · AI 翻译

在 @vllm_project 的 day-0 版本中,它带来了:

专用的 MSA 预填充/解码内核、支持前缀缓存和分块预填充的 1M 上下文服务,同时在 Hopper 和 Blackwell 上支持 BF16 和 MXFP8 🚀

这才是开放权重(open-weight)的正确做法。

感谢 @vllm_project、@NVIDIAAI、@AIatAMD、@inferact

引用vLLM@vllm_project
🎉 恭喜 @MiniMax_AI 发布 MiniMax M3!前沿的编程与智能体能力、原生图像和视频输入、计算机使用,以及 100 万 token 上下文窗口,全部集成在一个开放模型中。 M3 的核心是 MSA,一种全新的稀疏注意力架构:它不再对完整的 KV 缓存进行密集注意力计算,而是让每个查询对 128 token 的 KV 块打分,并仅对得分最高的块运行注意力。这正是让 100 万 token 上下文变得可实际服务的关键。 M3 在 vLLM 中实现首日支持,并已在 NVIDIA 和 AMD 硬件上验证: ✨ 带专用 prefill 和 decode 内核的 MSA 稀疏注意力 ✨ 支持前缀缓存和分块 prefill 的 100 万 token 上下文服务 ✨ BF16 和 MXFP8 检查点,同时为 Hopper 和 Blackwell 提供 MoE 后端 ✨ 原生多模态输入(图像 + 视频) ✨ 面向智能体工作负载的工具调用、推理解析和思考模式控制 这样的首日支持是真正的团队协作成果。感谢 @MiniMax_AI、@NVIDIAAI、@AIatAMD 和 @inferact 的团队,以及 vLLM 社区让这一切成为可能。🙏 深入了解实现细节、内核工作和部署方案: 🔗 https://vllm.ai/blog/2026-06-12-minimax-m3-vllm
原文

🎉 Congrats to @MiniMax_AI on releasing MiniMax M3! Frontier coding and agentic capabilities, native image and video input, computer use, and a 1M-token context window, all in a single open model. At the heart of M3 is MSA, a new sparse attention architecture: instead of attending densely over the full KV cache, each query scores 128-token KV blocks and runs attention only over the top blocks. That is what makes 1M-token context practical to serve. M3 runs in vLLM with day-0 support, verified on NVIDIA and AMD hardware: ✨ MSA sparse attention with dedicated prefill and decode kernels ✨ 1M-token context serving with prefix caching and chunked prefill ✨ BF16 and MXFP8 checkpoints, with MoE backends for both Hopper and Blackwell ✨ Native multimodal input (image + video) ✨ Tool calling, reasoning parsing, and thinking-mode control for agent workloads Day-0 support like this is a true team effort. Grateful to the teams at @MiniMax_AI, @NVIDIAAI, @AIatAMD, and @inferact, and to the vLLM community for making it happen. 🙏 Deep dive into the implementation, kernel work, and deployment recipes: 🔗 https://vllm.ai/blog/2026-06-12-minimax-m3-vllm

在 X 查看被引用的帖子

来源:MiniMax (official) · x.com