Qwen-RobotWorld:具身智能体的无界世界
Qwen-RobotWorld: Boundless Worlds for Embodied Agents
Qwen-RobotWorld以语言为统一动作接口,采用双流Multimodal Diffusion Transformer(MMDiT)架构,将Qwen2.5-VL作为动作编码器。在4个基准测试中取得顶尖成绩,统一20余种机器人形态,基于860万跨场景训练对和1300多项操作技能。语言接口标准化500多种动作类别,支持操作、自动驾驶、室内导航的联合训练。还支持Scene2Robot人类到机器人转移及2–4路多视角几何一致视频生成。
具身智能的世界模型长期受限于单一形态,Qwen-RobotWorld用语言统一动作接口,把操作、驾驶、导航合训,多视角几何一致性和人类演示迁移是过去一年最扎实的落地信号,做机器人的别错过。
Paper Embodied intelligence requires agents to perceive, reason about, and act within physical environments. World models offer a scalable path forward — but current approaches face a fundamental tension. General video generation models learn rich visual priors but lack the ability to model embodied physics. Domain-specific embodied models are tailored to individual scenarios and cannot generalize across embodiments.
Qwen-RobotWorld bridges this gap by treating natural language as a universal action interface. A single instruction like "pick up the red cup and place it on the shelf" implicitly encodes the complete action sequence, goal state, and physical constraints — no robot-specific control interface needed. This allows manipulation, autonomous driving, and indoor navigation to be trained jointly, with each domain's physical knowledge reinforcing the others.
Language unifies the action space: world knowledge and embodied knowledge reinforce each other within a single model, enabling cross-scenario, cross-task physical generalization.
Key Highlights#
Top-Tier Across 4 Benchmarks
20+Robot Embodiments Unified
8.6M Cross-Scenario Training Pairs
1300+Manipulation Skills
Language-Driven Unified Action Interface — natural language standardizes 20+ robot embodiments and 500+ action categories into one interface, enabling joint cross-scenario training Dual-Stream Diffusion World Model — MMDiT with Qwen2.5-VL as action encoder, combining deep language understanding with internalized physical world knowledge Cross-Scenario Physical Generalization — manipulation, driving, navigation, and human-to-robot transfer jointly trained under 8.6M video-text pairs Multi-View Geometrically Consistent Generation — synchronized 2–4 camera streams with 3D-consistent object identity and motion trajectories
Model Architecture#
Dual-Stream Diffusion World Model#
Qwen-RobotWorld adopts a dual-stream Multimodal Diffusion Transformer (MMDiT):
Understanding stream processes semantic features from a frozen Qwen2.5-VL encoder, representing the language action $a{t}$a t. Generation stream processes visual latents from a video-compatible VAE, representing the visual state $s{t}$s t.
The two streams interact via joint attention at every layer, enabling bidirectional cross-modal fusion throughout the denoising process.
Using an MLLM as the action encoder — rather than lightweight encoders like T5 or CLIP — provides two key advantages: (1) deep language understanding accurately parses complex, compositional instructions into precise condition signals; (2) internalized world knowledge (e.g., that robot arms are rigid bodies with fixed joint constraints) implicitly constrains physically plausible transitions, preventing common failure modes like object deformation across frames.
Scene2Robot: Human-to-Robot Transfer#
Scene2Robot enables cross-embodiment video editing: human demonstrations are retargeted to 14 robot morphologies via a multi-segment conditioning mechanism, where joint attention allows the generation to simultaneously attend to scene appearance and robot motion trajectory. This capability both serves as a data scaling engine during training and enables human-to-robot transfer at inference time.
Multi-View Geometrically Consistent Generation#
Single-camera observation inevitably occludes critical contact and spatial details. Qwen-RobotWorld generates 2–4 synchronized camera streams — main view, wrist-mounted views, and third-person views — with geometrically consistent object identity and motion across all viewpoints. During training, synchronized frames from multiple cameras are spatially concatenated into a single input; the model generates all views simultaneously, with asymmetric 3D RoPE providing spatial encoding and attention layers naturally establishing cross-view correspondence — without any architectural modification. This cross-view consistency further acts as a geometric regularizer, teaching the model object shape, depth, and spatial layout.
Data: Embodied World Knowledge#
EWK Dataset#
The Embodied World Knowledge (EWK) dataset is organized along four complementary axes, each targeting a distinct source of physical variation:
Multi-Embodiment — human hands, 7 robot arm configurations, ego vehicles, mobile agents, spanning 20+ distinct robot models Multi-Task — atomic manipulation skills, long-horizon compositions, locomotion, dynamic/deformable interactions across 500+ action categories Multi-Scenario — real-world first, sim-augmented: kitchens, workshops, outdoor settings, plus photorealistic simulation for downstream VLA evaluation Multi-View — main, wrist, and synchronized multi-view streams (~1.6M of 6M embodied samples include 2–4 view concatenations)
View Detailed Dataset Inventory
| Dataset | Embodiment | Views | Contribution |
|---|---|---|---|
| Manipulation (~5.9M samples) | |||
| EgoHOD, EPIC-Kitchens, Egocentric-10k | Human hands | Egocentric | Dexterity & coordination prior |
| Bridge V2, RH20T, Droid | Single-arm grippers | 3rd-person + wrist | Interaction primitives |
| Robomind, RoboCoin | Single/dual-arm, humanoids | Ego + 3rd-person | Cross-embodiment generalization |
| Agibot-World, Galaxea | Single-arm (gripper + dexterous) | Synced ego + wrist + 3rd | Temporal & multi-view consistency |
| Qwen-Aloha (internal) | Dual-arm grippers | Head + dual wrist | Multi-view grasping prior |
| ActionNet, OpenLoong | Dexterous hands | Wrist + 3rd-person | Fine-grained dexterity |
| Autonomous Driving (~200K samples) | |||
| Waymo, NVIDIA PhysicalAI-AD, Bench2Drive, Sekai | Ego vehicle | Surround-view | Large-scale ego-motion & multi-agent dynamics |
| Indoor Navigation (6K+ episodes) | |||
| VLNVerse | Mobile agent | Egocentric | Room-scale spatial reasoning |
| Human-to-Robot Transfer | |||
| Scene2Robot (synthesized) | 14 robot morphologies | Multi-view | Cross-embodiment video editing |
Action-Language Mapping#
The central challenge in building a universal world model is representational heterogeneity: manipulation uses joint angles, driving uses steering commands, navigation uses heading vectors — each requiring a separate model. Our action-language mapping framework resolves this by projecting all action signals onto a shared natural language space, so that videos from a Franka gripper, an autonomous vehicle, and a navigation agent all become instances of the same language-conditioned video generation task.
A hierarchical five-layer annotation pipeline ensures caption quality and precision:
Task GoalHigh-level intent — what should change between states
Action DetailSpatio-temporal trajectories with explicit viewpoint declaration
Physical FeedbackObservable consequences on the environment
Comprehensive CaptionFull description for precise prediction
Concise CaptionEssential elements for brief task-level commands
During training, comprehensive and concise descriptions are sampled with equal probability, so the model handles both detailed trajectory specifications and brief task-level commands.
Training#
Training follows a general-to-expert progressive curriculum:
| Stage | Phase | Data Mix | Objective |
|---|---|---|---|
| Pretraining | T2I / T2V / TI2V joint | General data | Build foundational visual priors |
| Human interaction | Ego4D, EPIC-Kitchen, etc. | Grasping & tool-use priors | |
| SFT | Phase 1: Single-view manipulation | Embodied + general joint training | Core manipulation physics |
| Phase 2: Multi-view expansion | Broaden viewpoint coverage | ||
| Phase 3: Multi-view concatenation | Cross-view geometric consistency | ||
| Phase 4: Complex cross-domain | Long-horizon & cross-scenario |
Pretraining on general data and human interaction videos (Ego4D, EPIC-Kitchen) builds broad visual priors — the T2I task specifically anchors object geometry that transfers to video generation through the shared backbone. SFT then progressively deepens embodied expertise across four phases while keeping general data in every batch, ensuring both capabilities advance together rather than trade off.
Demos#
Fine-Grained Language Grounding#
Given identical initial frames, the model produces qualitatively distinct videos when a single keyword differs. It also handles complex, multi-step instructions requiring long-horizon reasoning.
Contrastive Instruction Following:
L: Pick up the red strawberry
R: Pick up the yellow potato
L: Place pen on the wooden tray
R: Place pen on the white paper
L: Hand glue to the person
R: Place glue into the penholder
Complex Multi-Step Instructions:
Sequentially pick up the red and yellow bell peppers, place them on the table from left to right
Grab the yellow-and-blue stacked block, position it above the green-and-blue stacked block
Cross-Domain Generalization#
(A) Cross-Embodiment:
(B) Cross-Task × Cross-Environment:
(C) Multi-View Consistency:
(D) Zero-Shot Robustness:
Ours
LVP
Cosmos2.5-14B
Human-to-Robot Transfer#
The Scene2Robot mechanism preserves task intent from a human demonstration (left) while adapting motion to embodiment-specific kinematic constraints (right).
Beyond Manipulation: Driving and Navigation#
The learned world model generalizes beyond robot manipulation to broader mobility scenarios.
Autonomous Driving. Generated driving episodes from Bench2Drive, NVIDIA PhysicalAI-AD, Sekai, and Waymo demonstrate coherent scene dynamics and vehicle behaviors.
Indoor Navigation. Egocentric navigation episodes from VLNVerse show the model’s ability to simulate first-person movement through complex indoor environments.
From tabletop manipulation to autonomous driving and indoor navigation — Qwen-RobotWorld demonstrates that a unified world model can generalize beyond a single morphology or scenario family.
We evaluate against general video generation models (Sora2, Veo3, Wan2.6, Kling, LTX-2) and embodied world models (Cosmos, LVP, GigaWorld, Vidar, Wow) across four benchmarks.
EWMBench
4.60
Embodied motion fidelity
DreamGen
4.952
Instruction following & physics alignment
WorldModelBench
8.99
Physical reasoning & instruction following
PBench
0.804
Physical behavior evaluation
EWMBench
| Type | Model | SceneC | HSD | Dyn | nDTW | Diversity | BLEU | CLIP | Logics | Overall |
|---|---|---|---|---|---|---|---|---|---|---|
| General | Veo3 | 0.842 | 0.213 | 0.193 | 0.161 | 0.022 | 0.214 | 0.897 | 0.947 | 3.49 |
| Wan2.6 | 0.671 | 0.203 | 0.090 | 0.172 | 0.050 | 0.162 | 0.874 | 1.000 | 3.22 | |
| Kling | 0.821 | 0.327 | 0.182 | 0.342 | 0.017 | 0.259 | 0.901 | 1.000 | 3.85 | |
| LTX-2 | 0.785 | 0.208 | 0.128 | 0.244 | 0.012 | 0.143 | 0.887 | 0.500 | 2.91 | |
| Sora2 | 0.853 | 0.281 | 0.349 | 0.275 | 0.031 | 0.247 | 0.910 | 0.947 | 3.89 | |
| Embodied | Cosmos | 0.796 | 0.250 | 0.205 | 0.253 | 0.080 | 0.123 | 0.846 | 0.733 | 3.29 |
| GigaWorld | 0.871 | 0.305 | 0.085 | 0.278 | 0.028 | 0.205 | 0.887 | 0.900 | 3.56 | |
| LVP | 0.880 | 0.425 | 0.043 | 0.623 | 0.009 | 0.218 | 0.900 | 0.952 | 4.05 | |
| Vidar | 0.734 | 0.188 | 0.152 | 0.177 | 0.065 | 0.161 | 0.882 | 0.941 | 3.30 | |
| Wow | 0.887 | 0.249 | 0.053 | 0.257 | 0.027 | 0.193 | 0.900 | 0.952 | 3.52 | |
| Ours | 0.914 | 0.566 | 0.343 | 0.671 | 0.011 | 0.208 | 0.883 | 1.000 | 4.60 |
Strong motion fidelity (HSD 0.566), high scene consistency (0.914), and perfect logic constraint satisfaction.
DreamGen Bench
| Model | GR1-Env PA | GR1-Env IF | GR1-Object PA | GR1-Object IF | GR1-Behavior PA | GR1-Behavior IF | Total |
|---|---|---|---|---|---|---|---|
| Cosmos-sft | 0.709 | 0.655 | 0.775 | 0.720 | 0.649 | 0.621 | 4.129 |
| LVP | 0.810 | 0.772 | 0.745 | 0.829 | 0.713 | 0.889 | 4.758 |
| Vidar | 0.445 | 0.647 | 0.478 | 0.726 | 0.394 | 0.651 | 3.341 |
| GigaWorld | 0.621 | 0.933 | 0.500 | 0.852 | 0.426 | 0.884 | 4.216 |
| Wow | 0.793 | 0.826 | 0.755 | 0.849 | 0.809 | 0.696 | 4.728 |
| Ours | 0.828 | 0.793 | 0.840 | 0.878 | 0.781 | 0.832 | 4.952 |
Strong object-level compositional generalization (GR1-Object IF: 0.878), with consistent physics alignment across all subsets.
WorldModelBench
| Type | Model | Instr. (0-3) | Frame | Temp | Newton | Mass | Fluid | Penetr. | Grav. | Phys. | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| General | Veo3 | 2.52 | 0.98 | 0.95 | 1.00 | 0.89 | 0.99 | 0.91 | 1.00 | 4.80 | 9.25 |
| Wan2.6 | 2.50 | 0.99 | 0.95 | 1.00 | 0.89 | 0.99 | 0.94 | 1.00 | 4.83 | 9.27 | |
| Sora2 | 2.21 | 0.96 | 0.93 | 1.00 | 0.91 | 0.99 | 0.95 | 1.00 | 4.84 | 8.93 | |
| Kling | 1.59 | 0.97 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 5.00 | 8.55 | |
| LTX-2 | 1.97 | 0.69 | 0.62 | 0.99 | 0.60 | 1.00 | 0.73 | 1.00 | 4.32 | 7.61 | |
| Embodied | Cosmos | 2.14 | 1.00 | 0.94 | 1.00 | 0.92 | 1.00 | 0.94 | 1.00 | 4.86 | 8.94 |
| LVP | 2.01 | 0.89 | 0.91 | 1.00 | 0.93 | 0.99 | 0.95 | 1.00 | 4.87 | 8.67 | |
| GigaWorld | 2.13 | 0.59 | 0.46 | 1.00 | 0.48 | 0.99 | 0.69 | 0.98 | 4.13 | 7.31 | |
| Vidar | 1.62 | 0.54 | 0.45 | 1.00 | 0.56 | 1.00 | 0.85 | 1.00 | 4.40 | 7.01 | |
| Wow | 2.05 | 0.76 | 0.65 | 1.00 | 0.65 | 0.99 | 0.81 | 1.00 | 4.45 | 7.91 | |
| Ours | 2.33 | 0.87 | 0.85 | 1.00 | 1.00 | 1.00 | 0.94 | 1.00 | 4.94 | 8.99 |
Perfect physics adherence (1.00) across Newton's laws, mass conservation, fluid dynamics, and gravity, with strong instruction following (2.33/3.0).
PBench
| Type | Model | I2V-Bg | I2V-S | Aes | Img | Bg-Con | Mot | Sub-Con | O-Con | Quality | Domain | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| General | Veo3 | 0.975 | 0.980 | 0.526 | 0.698 | 0.938 | 0.994 | 0.927 | 0.128 | 0.771 | 0.882 | 0.827 |
| Wan2.6 | 0.856 | 0.843 | 0.514 | 0.719 | 0.906 | 0.978 | 0.843 | 0.136 | 0.724 | 0.832 | 0.778 | |
| Sora2 | 0.981 | 0.973 | 0.487 | 0.672 | 0.961 | 0.994 | 0.954 | 0.129 | 0.769 | 0.841 | 0.805 | |
| Kling | 0.982 | 0.979 | 0.521 | 0.699 | 0.920 | 0.990 | 0.927 | 0.124 | 0.768 | 0.874 | 0.821 | |
| LTX-2 | 0.948 | 0.955 | 0.506 | 0.622 | 0.932 | 0.986 | 0.904 | 0.118 | 0.746 | 0.845 | 0.796 | |
| Embodied | LVP | 0.979 | 0.981 | 0.515 | 0.679 | 0.954 | 0.991 | 0.962 | 0.116 | 0.772 | 0.812 | 0.792 |
| GigaWorld | 0.957 | 0.944 | 0.495 | 0.641 | 0.925 | 0.984 | 0.892 | 0.128 | 0.746 | 0.841 | 0.794 | |
| Wow | 0.967 | 0.957 | 0.517 | 0.689 | 0.941 | 0.980 | 0.929 | 0.111 | 0.761 | 0.786 | 0.774 | |
| Vidar | 0.935 | 0.922 | 0.501 | 0.573 | 0.912 | 0.982 | 0.863 | 0.120 | 0.726 | 0.810 | 0.768 | |
| Cosmos | 0.974 | 0.973 | 0.470 | 0.663 | 0.940 | 0.989 | 0.931 | 0.160 | 0.763 | 0.840 | 0.802 | |
| Ours | 0.956 | 0.943 | 0.455 | 0.649 | 0.956 | 0.990 | 0.933 | 0.124 | 0.751 | 0.857 | 0.804 |
Strong domain understanding (0.857) and motion smoothness (0.990), reflecting consistent temporal coherence across physical scenarios.
If you find our work helpful, feel free to give us a cite.
@article{qwenrobot-world, title={Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation}, author={Qwen Team}, year={2026}}
来源:Qwen:Blog Retrieval(API) · qwen.ai