跳到正文
北京时间
原文
NVIDIA Technical Blog:Agentic AI / Generative AI·· 2026-08-26精选AI 评分70

Qwen3.8-Flash-Next 发布并可在 NVIDIA GB300 NVL72 上运行智能体编码实验

Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding

AI 导读

阿里巴巴发布 Qwen3.8-Flash-Next 模型权重,作为即将到来的 Qwen4 架构预览。

推荐理由

原文给出 GDN 与 QSA 混合架构的设计细节和长上下文吞吐数据,可帮助读者评估该模型在智能体编码场景的部署价值。

正文

Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s a multimodal mixture-of-experts (MoE) model with a 125B-parameter main model supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. It has a native 262,144-token context window, extensible to 1M tokens with YaRN. 

NVIDIA provides best-effort Day 0 functional support through SGLang, vLLM, and NVIDIA TensorRT LLM, validation across NVIDIA GB300 NVL72 for inference, and post-training recipes from NVIDIA NeMo AutoModel and NVIDIA NeMo RL. 

Architectural innovations for long-context inference

Qwen3.8-Flash-Next is designed for high-volume, context-intensive applications such as agentic coding, document processing, and tool-driven workflows. As context grows, attention compute and KV cache memory become bottlenecks. The model addresses both with a hybrid architecture combining Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA). Three out of every four layers use GDN to continuously compress historical context into a fixed-size recurrent state, eliminating KV cache growth as sequences lengthen. The remaining layer uses QSA for precise retrieval across the full context. 

Previous sparse-attention approaches rely on token-level indexers that become increasingly computationally expensive as context length grows. QSA aggregates the sequence into micro-blocks, estimates their importance at the block level, and selects only the most relevant regions. This cuts attention, compute, and indexing overhead within each layer, making the design well-suited to architectures alternating between GDN and QSA layers. 

Alibaba’s published benchmarks suggest that QSA can improve the efficiency of 1M-token workloads. Compared with full attention, its attention kernel delivered speedups of up to 7.6x during prefill and 4.9x during decoding. In a cache-heavy online serving test at a 1M-token context length and with a 90% prefix-cache hit rate, Qwen3.8-Flash-Next achieved 8.6x the prefill throughput of Qwen3.7-Plus.

A diagram showing how GDN and QSA with MoE reduce memory and compute for large-context inference.
Figure 1. Overview of Qwen3.8-Flash-Next showing three layers of GDN and one layer of QSA with MoE to reduce memory and compute for large-context inference
视频 · 前往原文观看
Video 1. Qwen3.8-Flash-Next diagnoses and fixes a bug

Running Qwen3.8-Flash-Next on NVIDIA GB300 NVL72

The GB300 NVL72 features a rack-scale architecture that integrates 72 NVIDIA Blackwell Ultra GPUs into a single platform. Its large, 72-GPU NVIDIA NVLink domain enables efficient all-to-all communication at 130 TB/s, eliminating bottlenecks that appear when expert traffic must cross traditional off-the-shelf networks. Running on NVIDIA GB300 NVL72 delivers over16K tokens per second per GPU and over 200 tokens per second per user, enabling developers to experiment with agentic coding applications at high throughput and low latency. 

A chart showing Qwen3.8-Flash-Next FP8 performance on NVIDIA GB300 NVL72 throughput vs. interactivity using TensorRT -LLM.
Figure 2. A Pareto curve showing Qwen3.8-Flash-Next achieving peak throughput above 16K tokens per second per GPU on NVIDIA GB300 NVL72 

Beyond rack-scale deployment, Qwen3.8-Flash-Next also runs on local NVIDIA hardware, including NVIDIA DGX Station, NVIDIA DGX Spark clusters, and workstations with four NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition GPUs. Developers can prototype and evaluate agentic coding workflows on local hardware and scale the same model to GB300 NVL72 for production serving. 

Post-train Qwen3.8-Flash-Next and serve it with your preferred inference engine

Developers can fine-tune the model for domain-specific use cases using NVIDIA NeMo AutoModel, a PyTorch-native fine-tuning library with Day-0 Hugging Face checkpoint support. Train directly on existing checkpoints without model conversion, with support for full SFT or memory-efficient LoRA fine-tuning. Users can go a step to perform reinforcement learning using NVIDIA NeMo RL recipes. 

NVIDIA supports multiple inference stacks to meet a variety of developer needs. SGLang, vLLM, and TokenSpeed provide open-source inference recipes for developers requiring greater control over performance on the NVIDIA-accelerated platform.  

Get started with Qwen3.8-Flash-Next

Try the model directly from QwenCloud.  

Download the model weights from Hugging Face or ModelScope and deploy with a model-free NVIDIA NIM from NVIDIA NGC. 

来源:NVIDIA Technical Blog:Agentic AI / Generative AI · developer.nvidia.com