跳到正文
北京时间
原文
LMSYS:Blog(Chatbot Arena 团队)· AMD & SGLang Team·· 2026-05-28精选AI 评分69

SGLang 团队与 AMD 合作,使 AMD Instinct™ MI355X GPU 的大规模 DeepSeek-R1 分离式推理在总拥有成本上具备竞争力

Blog Win on TCO: How AMD Instinct™ MI355X Achieves Cost-Competitive Distributed Inference Through SGLang with MoRI The SGLang and AMD team has worked closely to unlock competitive Total Cost of Ownership (TCO) for large-scale DeepSeek-R1 disaggregated inference on AMD Instinct™ MI355X GPUs. Building on SGLang's se... AMD & SGLang Team

AI 导读

SGLang 与 AMD 团队合作,通过一系列全栈优化,使 AMD Instinct™ MI355X GPU 在运行 DeepSeek-R1 大模型推理时实现了极具竞争力的总拥有成本。在 129 tok/s/user 的交互延迟下,其成本为每百万 token $0.169,比 NVIDIA B200(Dynamo TRT-LLM)方案低 5%,比 B200(SGLang)方案低 40%。吞吐量方面,24 块 AMD GPU 达到 2,436 tok/s/GPU,比使用 48 块 GPU 的 B200 SGLang 方案每 GPU 吞吐量高 1.25 倍。核心优化包括:MoRI 混合 FP4/FP8 量化全到全通信、MoRI-IO KV Cache 后端、两批重叠与 SDMA、ROCm 上的 Specv2 MTP 以及 CPU 流式处理优化。

推荐理由

AMD MI355X跑DeepSeek-R1的TCO比NVIDIA B200低5%,吞吐还高出1.25倍,这是开源框架SGLang对闭源生态的一次真实挑战,做推理部署的应该点开看看完整的全栈优化。

正文

‹ Back to Blog Contents Results at a Glance Key Optimizations MoRI Quantized All-to-All for Expert Parallelism Hybrid FP4/FP8 quantized all-to-all Adaptive kernel selection MoRI-IO KV Cache Backend Inline transfer for high-concurrency KV migration Broader model coverage Two-Batch Overlap (TBO) with SDMA FlyDSL FusedMoE for High-Performance MoE Compute Specv2 MTP on ROCm CPU Streaming Optimization Looking Ahead Summary Acknowledgements References Endnotes Win on TCO: How AMD Instinct™ MI355X Achieves Cost-Competitive Distributed Inference Through SGLang with MoRI

AMD & SGLang Team May 28, 2026

The SGLang and AMD team has worked closely to unlock competitive Total Cost of Ownership (TCO) for large-scale DeepSeek-R1 disaggregated inference on AMD Instinct™ MI355X GPUs. Building on SGLang's serving framework and AMD's MoRI communication library, we demonstrate that AMD achieves competitive — and at key operating points, superior — TCO compared to NVIDIA B200 running Dynamo + TRT-LLM. These results are validated by InferenceX, SemiAnalysis's open-source continuous benchmark platform that tests across hundreds of GPUs with a live dashboard.

This post describes what we achieve, how we achieve it, and our plans for the road ahead.

TL;DR At 129 tok/s/user interactivity, AMD Instinct™ MI355X delivers $0.169 per million tokens — 5% lower cost than B200 TRT-LLM and 40% lower cost than B200 SGLang. 2,436 tok/s/GPU on 24 GPUs — 1.25× higher throughput per GPU than B200 SGLang (48 GPUs). Full-stack optimizations: AITER GEMM tuning, MoRI quantized all-to-all (up to 2.56× bandwidth reduction), MoRI-IO KV cache backend (~10% higher throughput than Mooncake), Two-Batch Overlap with SDMA, Specv2 MTP on ROCm, and CPU streaming optimization.

Figure 1: InferenceX TCO comparison — AMD Instinct™ MI355X MoRI SGLang vs B200 Dynamo SGLang. Full TCO breakdown in SemiAnalysis TCO model, or call Andrew Lekashman at +1 (408) 404 8069 Results at a Glance

At the typical operating point representative of production coding assistants and interactive chatbots — e.g. 129 tok/s/user interactivity — we observe the following: AMD Instinct™ MI355X (MoRI SGLang MTP): $0.169 per million tokens, 2,436 tok/s/GPU (24 GPUs) NVIDIA B200 (Dynamo TRT-LLM MTP): $0.178 per million tokens, 3,128 tok/s/GPU (28 GPUs) NVIDIA B200 (Dynamo SGLang MTP): $0.284 per million tokens, 1,945 tok/s/GPU (48 GPUs)

AMD Instinct™ MI355X delivers 5% lower cost than B200 TRT-LLM, 40% lower cost than B200 SGLang, and 1.25× higher throughput per GPU than B200 SGLang — winning on both cost and performance simultaneously.

Figure 2: Full pareto curve — throughput vs interactivity for AMD Instinct™ MI355X and B200 configurations Key Optimizations

We achieved these results through a series of full-stack optimizations spanning communication, compute kernels and serving infrastructure. The following sections walk through each in detail. MoRI Quantized All-to-All for Expert Parallelism Hybrid FP4/FP8 quantized all-to-all

We built hybrid quantized all-to-all in a series of PRs: FP4 dispatch + FP8 combine direct cast (#19757) introduced MXFP4 dispatch to reduce communication latency; FP8 blockwise combine (#24879) added fine-grained FP8 blockwise quantization for the combine path, achieving ~2% higher accuracy than direct-cast FP8; and auto-select dispatch dtype (#21040) made MoRI automatically detect the correct dispatch quantization type from the model's MoE weight dtype, eliminating manual env var configuration.

In expert-parallel MoE inference, each token must be dispatched to the top-k selected experts via dispatch and combine communication primitives. For DeepSeek-R1 with a hidden dimension of 7,168 and top-8 expert routing, BF16 communication volume is significantly higher than that of FP8 and FP4 quantized communication.

The key insight is that on-the-fly MXFP4 quantization of dispatch will bring faster transmission with accuracy lossless. Similarly, expert outputs (combine phase) tolerate FP8 quantization without meaningful accuracy loss.

MoRI supports multi-level quantized communication:

MoRI-EP combine kernel micro-benchmark on AMD Instinct™ MI355X (EP8, BF16 input, max_tokens=4096, hidden_dim=7168, scale_dim=56, zero-copy=0, dispatch=128/16, combine=128/16, 10-round average, combine latency only):

Case Path Combine Latency
Normal (no-scale) fp8_blockwise specialized ~736 µs
Uniform[−1024, 1024] (scale-active) fp8_blockwise specialized ~770 µs
Force-scale-active fp8_blockwise specialized ~769 µs
Reference bf16 no-quant ~907 µs

For MXFP4 models such as amd/DeepSeek-R1-0528-MXFP4-v2, the system uses FP4 dispatch + FP8 combine, achieving a 2.56× overall round-trip bandwidth reduction (from 28,672 to 11,200 bytes per token).

Blockwise quantization preserves accuracy through fine-grained scaling. By default, FP8 blockwise uses per-128-element FP32 scale factors, achieving a good tradeoff between performance and accuracy.

The quantization mode is auto-detected from the model's weight format and can be overridden via SGLANG_MORI_DISPATCH_DTYPE and SGLANG_MORI_COMBINE_DTYPE environment variables. Adaptive kernel selection

We added inter-node kernel type switching (#18437) to MoRI-EP — the new InterNodeV1LL kernel delivers 1.52× dispatch and 1.82× combine speedup over the original InterNodeV1 when the number of tokens per rank is below 256.

MoRI dynamically selects the optimal communication kernel based on workload characteristics:

Kernel Condition Optimized For
IntraNode Single-node (≤8 GPUs) Shared memory / P2P
InterNodeV1 Multi-node, >256 tokens/rank High throughput, staged RDMA
InterNodeV1LL Multi-node, ≤256 tokens/rank Low latency
AsyncLL SDMA-enabled paths Fully async send/recv split

The switching threshold is automatically configured based on the decode batch size, ensuring that prefill phases (large batches) use high-throughput kernels while decode phases (smaller per-rank batches) use low-latency kernels. MoRI-IO KV Cache Backend

In #22665 we overhauled MORI-IO with state transfer support (Mamba, SWA, NSA), a lock-free inline transfer model that eliminates worker-thread dispatch, and high-concurrency fixes for robust operation under thousands of concurrent requests. Inline transfer for high-concurrency KV migration Lock-free inline execution. Transfer requests execute directly in the caller path instead of being dispatched to worker threads. Transfer plans are precomputed once and reused across all layers, eliminating per-layer scheduling overhead and reducing lock contention. Robust at scale. Default RDMA parallelism is increased to 4 queue pairs and 4 workers per transfer, with thread-safe connection reuse that prevents port exhaustion under thousands of concurrent requests. Broader model coverage

Beyond standard MLA-based KV cache, MoRI adds state transfer support for hybrid architectures — Mamba (SSM state), SWA, and NSA — enabling disaggregated serving for models like Qwen3.5-397B-A17B. It also handles TP-mismatch scenarios where prefill and decode use different tensor-parallel degrees, correctly mapping replicated attention heads across ranks.

MoRI-IO benchmark on AMD Instinct™ MI355X (8 GPUs/node, 8× AMD Pensando Pollara 400 AI-NIC, DeepSeek-R1 671B FP8, TP=8, 2048 prompts, ISL=8192, OSL=1024):

Metric MoRI-IO Mooncake
Request throughput 7.49 req/s 6.80 req/s
Input token throughput 31,111 tok/s 28,257 tok/s
Output token throughput 3,775 tok/s 3,428 tok/s
Total token throughput 34,886 tok/s 31,685 tok/s

MoRI-IO delivers ~10% higher throughput than Mooncake across all metrics, with comparable single-request latency (~7 ms TPOT) and high accuracy (GSM8K 5-shot: 0.970). Two-Batch Overlap (TBO) with SDMA

TBO was built across three PRs: two-batch overlapping for MoRI EP (#19216) introduced the core dual-stream pipeline with MoRI's async API, delivering up to +25% throughput at large batch sizes; SDMA path for MoRI EP (#23929) enabled AMD's System DMA engines for zero-compute-overhead data movement via split send/recv; and dual-stream MoE on ROCm (#24005) activated the shared-expert overlap stream on the ROCm path, reducing mean TPOT from 97 ms to 83 ms.

Even with 2–4× bandwidth reduction from quantization, all-to-all communication remains significant. Two-Batch Overlap (TBO) hides this latency by interleaving communication and compute across two micro-batches:

  1. MicroBatch A dispatch sends quantized tokens over the network on a dedicated communication stream
  2. While network transfer is in flight, MicroBatch B attention computes on the main compute stream
  3. MicroBatch A arrives; MoE GEMM runs
  4. MicroBatch A combine sends results back on the communication stream
  5. Meanwhile, MicroBatch B dispatch begins

The dispatch and combine operations are split into A/B phases — dispatch_a for local quantization on the compute stream, dispatch_b for network transfer on the communication stream. A CommStreamPool manages dedicated streams, and events synchronize handoff points.

When SDMA is enabled (MORI_ENABLE_SDMA=true), data transfers run on AMD's dedicated System DMA engines that move data between GPU memory and network interfaces without consuming any compute units. This achieves true zero-compute-overhead communication, keeping every compute unit available for GEMM operations throughout the pipeline.

Figure 3: Two-Batch Overlap pipeline diagram — interleaved compute and communication streams FlyDSL FusedMoE for High-Performance MoE Compute

Traditionally, FusedMoE kernels on AMD relied solely on Composable Kernel (CK) — hand-tuned templates that are performant but inflexible. AITER introduces FlyDSL (Flexible Layout Python DSL), a Python DSL backed by an MLIR stack for authoring GPU kernels with explicit layouts and tiling, as a competitive FusedMoE kernel path for mixed-precision MoE (e.g., A4W4) on MI355X. FlyDSL enables rapid exploration of kernel configurations beyond what hand-tuned CK templates cover, and at a typical concurrency of 512, we gained up to 1.6× latency reduction for the FusedMoE compute.

MoE GEMM performance is shape-dependent, and the dominant shapes differ by serving scenario. In low-latency pure TP deployments, each GPU processes all experts with small batch sizes, producing tall-skinny GEMMs. In high-throughput DP+EP deployments, tokens are distributed across expert-parallel ranks, yielding different N/K dimensions per expert. FlyDSL allows us to provide separate tuning configurations for each scenario to maximize MI355X utilization.

Triton blockscale GEMM tuning — alongside FlyDSL, the A8W8 blockscale GEMM path uses per-shape tuned configurations for MI355X (gfx950). Key shapes like (N=7168, K=16384) and (N=16384, K=1536) — matching DeepSeek-R1's expert dimensions — are tuned with optimized block sizes, warp counts, pipeline stages, and k-splitting parameters. Special-case tuning for ultra-small M values (≤8, ≤256) targets the small per-expert batches typical in EP decode.

_Figure 4: FlyDSL kernel and Tri…

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org