跳到正文
北京时间
原文
Qwen:Blog Retrieval(API)· QwenTeam·· 2026-06-16精选AI 评分72

Qwen-RobotManip:对齐解锁机器人操作基础模型的规模化能力

Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

AI 导读

Qwen-RobotManip 是通义千问基于 Qwen-VL 的视觉-语言-动作(VLA)基础模型,引入覆盖表示、运动和行为三维度的统一对齐框架。仅使用开源机器人数据集和人演示视频,构建约 38,100 小时预训练语料,涵盖 15 种机器人形态。在 LIBERO-Plus 达 91.4%,RoboTwin-C2R Hard 达 69.4%,RoboCasa365 Composite-Unseen 达 14.9%,EBench 达 45.6%,RoboTwin-IF 达 72.0%,并在 RoboChallenge Table30 v1 generalist track 夺冠。模型采用 80 维状态-动作表示、人-机器人数据合成管道(1,933 小时第一人称视频转 24,808 小时数据)及上下文策略适配。

推荐理由

Qwen 这次发布的机器人模型,用统一对齐框架把跨实体数据规模化训练跑通了,OOD 泛化大幅领先,做具身智能的值得认真看一下。

正文 · 原文

Qwen-Omni × Qwen-RobotManip — Qwen-Omni observes the scene, randomly proposes manipulation tasks via speech, and judges execution in real time. Each video shows Qwen-RobotManip completing tasks on the fly with no pre-defined task list, demonstrating open-ended instruction following and generalization.

Qwen-RobotManip is validated across various real-robot platforms and tasks, demonstrating strong generalization to novel scenes, unseen language instructions, and cross-embodiment transfer.

View Real-World Evaluation Gallery

Foundation models in language and multimodality have achieved remarkable generalization because heterogeneous data sources can be aligned under a unified formulation, and abundant low-cost internet data allows diverse training signals to reinforce one another at scale. But can this scaling recipe be applied to robotic manipulation?

This is challenging. Unlike text or images, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity. Aligning representations across different robot embodiments, sensors, and task domains while simultaneously scaling the data has remained an open problem.

Qwen-RobotManip is a generalizable Vision-Language-Action (VLA) foundation model built upon Qwen-VL. It introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. Using only open-source robotic manipulation datasets and human demonstration videos without any proprietary data collection, Qwen-RobotManip constructs a ~38,100 hours pretraining corpus and already exhibits emergent generalization capabilities.

Without unified cross-embodiment alignment, scaling data produces conflicts; without data diversity, alignment alone cannot generalize. Alignment and scale are tightly coupled prerequisites for robotic foundation models.

Key Highlights#

Alignment

Representation · Motion · Behavior

Three-Dimensional Alignment

Open-Source Data Only

38K Hours of Manipulation Data

Across 15 Embodiments

Dominant

OOD Generalization

Across All Benchmarks

#1

RoboChallenge Table30 v1 Generalist Track

Sweeping Top 2, 20% Ahead of 3rd Place

Unified Cross-Embodiment Alignment Framework — a unified 80-dimensional state-action representation accommodates diverse embodiments, camera-frame end-effector delta poses make visually similar motions numerically proximate, and in-context policy adaptation reads execution history as an implicit embodiment identifier — together enabling consistent signal extraction across embodiments Human-to-Robot Synthesis at Scale — a pipeline converting 1,933h of egocentric human video into 24,808h of robot demonstrations across 15 embodiments via action retargeting, hand removal and inpainting, simulated rendering, and depth-guided compositing, coupled with a multi-stage curation pipeline ensuring data quality OOD Generalization: LIBERO-Plus 91.4% (+7.0 over π0.5), RoboTwin-C2R Hard 69.4% (+21.5 over π0.5), RoboCasa365 Composite-Unseen 14.9% (3× next best), EBench 45.6% (+18.5 over next best); RoboTwin-IF 72.0% (+22.4 over π0.5) confirming genuine language-conditioned control; 3× next best on RoboTwin-XE showing zero-shot cross-embodiment transfer Strong Real-World Performance: #1 on RoboChallenge Table30 v1 generalist track with 45% SR, sweeping top 2 and leading 3rd place by 20%; validated on real-robot platforms with 2× prior SOTA on in-domain and OOD tasks, few-shot adaptation, and cross-embodiment skill transfer

Scaling Manipulation Data#

Human-to-Robot Data Synthesis#

Robot manipulation data is scarce and expensive to collect. We introduce a Human-to-Robot synthesis pipeline that converts egocentric human manipulation videos into robot demonstrations across 15 robot embodiments via human-to-robot retargeting, hand removal and inpainting, and depth-guided robot compositing.

Data Sources#

The resulting pretraining corpus totals over 38,100 hours from three complementary sources:

Robot data (~11,420h): Open-source robotic datasets covering single-arm, dual-arm, and mobile manipulation.

Egocentric human data (~1,933h): Human manipulation videos collected from open-world environments, providing rich object-interaction and scene priors.

Human-to-Robot synthesized data (~24,808h): Generated from the egocentric data above across 15 robot platforms, serving as the primary scaling engine.

Data Curation#

We design a multi-stage curation pipeline to ensure VLA training data quality and annotation correctness. Five state-action filtering stages remove noisy actions, fix temporal misalignment, and verify kinematic consistency. Three cross-modal checks then validate that language instructions match the video content, that visual observations agree with recorded robot states, and that video frames are free of corruption.

Qwen-RobotManip Model Design#

Qwen-RobotManip couples a Qwen3.5-4B vision-language backbone with a flow-matching Diffusion Transformer (DiT) action head. Three design choices enable coherent cross-embodiment training:

Canonical State-Action Representation. All robot states and actions are mapped to a unified 80-dimensional vector covering single-arm, dual-arm, dexterous hand, and mobile base configurations. A per-dimension binary mask ensures gradients flow only through populated slots, ensuring different embodiments share the same representation without conflict.

Camera-Frame Delta Pose. End-effector actions are expressed as deltas in the camera coordinate frame rather than the robot base frame, making visually similar actions numerically proximate across embodiments. Camera extrinsics are injected via Camera Positional Encoding (CaPE) in the cross-attention layers, while intrinsics are encoded into visual tokens for field-of-view awareness. The DiT is further conditioned on end-effector type embeddings for embodiment-aware action denoising.

In-Context Policy Adaptation. The model conditions action prediction on a structured embodiment prompt (specifying robot platform, execution speed, and FPS) together with a historical observation-action chunk, enabling on-the-fly adaptation to different embodiments and behavior patterns. A stochastic context sampling strategy during training prevents action-copy shortcuts and forces genuine policy learning.

Training. Pre-training uses dual-stream co-training with a VLA stream (robot manipulation data) and a VLM stream (vision-language understanding data) at a 9:1 ratio. Post-training adopts generalist SFT on all demonstration data collected for each benchmark. We propose co-training with VL data and VLA data during post-training, which further improves OOD instruction following and generalization.

Evaluation#

Qwen-RobotManip is evaluated across 500+ simulation tasks and 80+ real-world tasks spanning various robot embodiments.

Why OOD Evaluation Matters#

A critical finding in our experiments: standard benchmarks systematically fail to capture the quality of pretraining. On in-distribution benchmarks like LIBERO and RoboTwin, models trained from scratch without any large-scale robot pretraining achieve performance comparable to previous SOTA pretrained models. Strong IID scores do not indicate genuine generalization; they can be achieved through pattern matching alone.

The separation only becomes visible under out-of-distribution evaluation: novel scenes and task variations, following unseen instructions, and cross-embodiment transfer. This is why Qwen-RobotManip adopts OOD benchmarks as the north star for evaluating robotic foundation models.

In-Distribution Results#

On standard benchmarks, Qwen-RobotManip matches or exceeds previous SOTA.

Model LIBERO RT-Easy RT-Hard
$\pi{0}$π 0​ 94.4 65.9 58.4
$\pi{0.5}$π 0.5​ 97.6 82.7 76.8
StarVLA 98.0 85.7 87.3
Abot-M0 98.6 86.1 85.1
Being-H0.7 99.2 90.2 89.6
Qwen-RobotManip-scratch 98.2 88.7 88.4
Qwen-RobotManip 99.1 93.4 92.5
Qwen-RobotManip-Context 99.2 93.7 94.0

Out-of-Distribution Generalization#

Qwen-RobotManip substantially outperforms all previous models across three OOD generalization axes: task and scene variations, instruction following, and cross-embodiment transfer.

Detailed Analysis per Benchmark LIBERO-Plus — OOD robustness evaluation under 7 perturbation dimensions (camera, robot, language, lighting, background, noise, layout):

Model Camera Robot Language Light Background Noise Layout Total
$\pi{0}$π 0​ 13.8 6.0 58.8 85.0 81.4 79.0 68.9 53.6
$\pi{0.5}$π 0.5​ 78.4 73.6 80.8 96.2 94.1 89.0 84.5 84.4
StarVLA 52.5 49.8 88.5 95.7 95.7 73.0 76.9 74.1
Abot-M0 60.4 67.9 86.4 96.2 91.6 86.4 82.6 80.5
Being-H0.7 82.0 59.0 82.8 97.8 90.0 93.5 88.5 84.8
Qwen-RobotManip 87.2 75.5 85.6 96.6 97.7 97.7 87.3 89.0
Qwen-RobotManip-Context 89.9 83.9 86.5 98.6 99.9 97.9 87.5 91.4

Qwen-RobotManip achieves 89.0% and Qwen-RobotManip-Context achieves 91.4% overall. The per-dimension breakdown shows that Camera and Robot perturbations benefit most from large-scale robot data pretraining, while Language and Light robustness is already provided by the VLM backbone.

RoboTwin-Clean2Rand — Models are fine-tuned on the Clean dataset and tested under progressive environmental randomizations:

Model Easy Background Light Clutter Height Hard
StarVLA 58.1 27.1 50.9 24.2 48.4 10.6
GR00T-N1.7 43.6 40.4 41.9 27.1 39.0 20.7
$\pi{0.5}$π 0.5​ 73.1 67.0 69.2 57.9 67.6 47.9
Qwen-RobotManip 73.2 74.6 68.4 61.3 71.0 62.6
Qwen-RobotManip-Context 84.7 82.4 84.2 75.4 79.5 69.4

Qwen-RobotManip achieves the highest Hard success rate (62.6%), retaining ~86% of its Easy performance compared to 66% for $\pi{0.5}$π 0.5​ and under 30% for models without pretraining. Qwen-RobotManip-Context further achieves 69.4% on Hard, demonstrating the effectiveness of in-context policy adaptation.

RoboCasa365 — Evaluation across atomic and long-horizon manipulation in diverse kitchen environments:

Model Atomic Composite-Seen Composite-Unseen Total
$\pi{0}$π 0​ 36.3 5.2 0.7 15.0
$\pi{0.5}$π 0.5​ 39.6 7.1 1.2 16.9
GR00T-N1.5 50.7 14.8 2.7 23.9
RLDX-1 63.0 27.5 5.4 33.2
Qwen-RobotManip 68.6 20.1 14.9 35.9
Qwen-RobotManip-Context 63.9 22.6 11.2 33.8

On Composite-Unseen, which requires completing long-horizon tasks in OOD scenes, Qwen-RobotManip achieves 14.9%, nearly 3× the next-best model (5.4%).

EBench — Mobile manipulation across tabletop, pick-and-place, and long-horizon tasks:

Model TableTop SimplePnP LongHorizon Overall
SR Score SR Score SR
π₀ 15.7 30 35.0 39
π₀.₅ 12.9 32 45.0 50
X-VLA 8.6 24 50.0 54
InternVLA-A1 4.3 11 43.0 47
Qwen-RobotManip 50.0 70 56.5 60
Qwen-RobotManip-Context 49.3 56 55.0 66

Qwen-RobotManip achieves 45.6% overall SR and a composite score of 60, outperforming $\pi{0.5}$π 0.5​ (27.1% / 41) and all other baselines by a large margin across every split.

RoboTwin-IF — Instruction following with held-out unseen instruction templates:

Model Pick-Diverse Place-Rel. Ope.-Mic-Dr. Ope.-Stapler Ope.-Table Average
StarVLA 11 13 0 49 74 29.4
GR00T-N1.7 20 17 0 14 32 16.6
$\pi{0.5}$π 0.5​ 44 20 15 92 66 49.6
Qwen-RobotManip 79 57 42 90 93 72.2
Qwen-RobotManip-Context 77 71 33 89 90 72.0

Qwen-RobotManip achieves 72.2% average, a +22.6 point lead over $\pi{0.5}$π 0.5​. The largest gains appear on tasks where the instruction must be parsed to select the correct action among multiple plausible alternatives, confirming genuine language-conditioned control.

Zero-Shot Cross-Embodiment — Trained on AgileX ALOHA data only, evaluated on unseen robot embodiments on our RoboTwin-XE benchmark:

Model ARX UR5 Franka Total
$\pi{0.5}$π 0.5​ (joint) 24.6 2.2 0.9 9.2
$\pi{0.5}$π 0.5​ (eef) 11.5 10.0 1.1 7.5
Qwen-RobotManip (joint) 37.6 4.1 1.8 14.5
Qwen-RobotManip (eef) 42.9 22.8 5.9 23.9

Camera-frame EEF representation dramatically improves zero-shot transfer by abstracting away morphological differences. Qwen-RobotManip (eef) reaches 23.9% overall, 3.2× $\pi{0.5}$π 0.5​ (eef) at 7.5%.

Data Scaling: Alignment Enables Scale#

An important finding: only models with unified cross-embodiment representations exhibit clean log-linear data scaling behavior. Without the alignment framework (UnifiedSpace + UnifiedEEF), adding more data produces erratic or flat scaling curves. This confirms that alignment is the prerequisite for scale, not the other way around.

Real-World Experiments#

Generalization to Novel Scenes and Instructions#

In-domain evaluation across 7 tasks spanning basic pick-and-place, deformable object handling, and precision assembly:

Task $\pi{0.5}$π 0.5​ StarVLA Ours
table-cleanup 4/5 0/5 5/5
three-bowl-stacking 5/5 4/5 5/5
melon-in-bowl 2/5 0/5 5/5
towel-folding 4/5 3/5 4/5
block-in-drawer 0/5 0/5 5/5
yellow-disc-insertion 0/5 0/5 2/5
three-block-stacking 0/5 0/5 5/5
Average 42.9% 20.0% 88.6%

Out-of-domain evaluation with distribution shifts in visual scenes, objects, and instructions:

Task OOD Factors $\pi{0.5}$π 0.5​ StarVLA Ours
target-object-in-basket cluttered bg, unseen objects 8/10 0/10 10/10
left-right-bowl-stacking cluttered bg, left-right reference 1/10 0/10 10/10
tool-on-towel unseen small objects, distractors 0/10 0/10 6/10
banana-on-towel dynamic lighting (disco light) 6/10 0/10 9/10
Average 37.5% 0.0% 87.5%

Qwen-RobotManip achieves 88.6% in-domain and 87.5% OOD success, substantially outperforming $\pi{0.5}$π 0.5​ (42.9% / 37.5%) and StarVLA (20.0% / 0.0%).

Data-Efficient Skill Transfer Across Embodiments#

Few-shot adaptation. All methods are jointly finetuned on only 130 teleoperated demonstrations across 5 tasks. Qwen-RobotManip outperforms both baselines on 4 of 5 tasks:

Task Sub-step StarVLA $\pi{0.5}$π 0.5​ Ours
Put Fruits Place 1 / 2 / 3 3/1/0 9/5/2 9/5/3
Avg. success 13.3% 53.3% 56.7%
Put Blocks Open / Place1 / Place2 / Close 1/1/0/0 4/2/2/2 5/4/3/3
Avg. success 5.0% 25.0% 37.5%
Fold Towel Fold 1 / Fold 2 0/0 3/1 3/3
Avg. success 0.0% 20.0% 30.0%
Insert Screw Handover / Insert 0/0 2/0 2/0
Avg. success 0.0% 10.0% 10.0%
Unscrew Cap Grasp / Unscrew / Place 4/0/0 9/2/1 9/4/3
Avg. success 13.3% 40.0% 53.3%

Cross-embodiment skill transfer. A single policy jointly finetuned on 6K CobotMagic and 130 ARX demonstrations is evaluated on 4 novel tasks on ARX, for which ARX has zero training demonstrations:

Model Stack Plates Stack Blocks Fruits in Plate Trash in Bucket Avg.
w/o UnifiedSpace 0/10 0/10 3/10 0/10 7.5%
w/o UnifiedEEF 0/10 0/10 5/10 0/10 12.5%
Qwen-RobotManip 3/10 5/10 7/10 7/10 55.0%

The full unified framework achieves 55.0%, over 4× the best ablation variant, demonstrating that the unified representation enables skill-level transfer across kinematically different embodiments.

Complex Multi-Step Tasks and Emergent Recovery#

In the RoboChallenge Table30 v1 Generalist Track, which spans 30 tasks across 4 robot platforms, Qwen-RobotManip ranks 1st with 45% success rate and 59.83 process score, outperforming the runner-up by 20%. On 8 bimanual coordination tasks, it achieves 40% success versus $\pi{0.5}$π 0.5​’s 21.2%.

Bimanual coordination. Among the 30 benchmark tasks, 8 require tight bimanual coordination on the ALOHA platform, where the two arms must jointly stabilize, transport, and manipulate objects. Qwen-RobotManip achieves 40% average success rate, far exceeding $\pi{0.5}$π 0.5​ (21.2%), DM0 (16.2%), GR00T-MULTI (7.5%), and $\pi{0}$π 0​ (7.5%). Notably, Qwen-RobotManip is the only model to succeed on “pour fries into plate” (30% vs. 0% for all baselines), a task demanding sequential bimanual steps — stabilizing the fries box with the left arm, opening it with the right arm, picking it up, and pouring the contents onto the plate. We attribute this strong bimanual performance to two factors: (1) our pretraining corpus contains a substantial proportion of bimanual demonstration data, enabling the model to learn coordinated dual-arm control primitives; and (2) the Human-to-Robot synthesis pipeline further expands the effective bimanual pretraining data by synthesizing bimanual robot demonstrations from egocentric human videos.

Robust pick-and-place across embodiments. We identify 12 tasks across all four platforms that center on pick-and-place primitives, ranging from single-object grasping to multi-step sequential manipulation involving 4–5 objects. Qwen-RobotManip achieves 63.3% average success rate on these tasks, surpassing the next-best baseline DM0 (48.3%) by 15.0 percentage points. We attribute this capability to two factors: (1) the large-scale cross-embodiment pretraining data encodes abundant pick-and-place patterns, and (2) the unified action space enables knowledge sharing of fundamental spatial skills across different robot morphologies.

Reactive error recovery. When objects slip during grasping, the model autonomously retries until successful. This behavior emerges from pretraining at scale, not from explicit programming.

Towards Scalable Robotic Foundation Models#

Qwen-RobotManip demonstrates that the scaling recipe behind language and multimodal foundation models can be extended to robotic manipulation, but only when alignment and scale work in tandem. The unified cross-embodiment representation is what makes large-scale multi-source training productive rather than conflicting, and the Human-to-Robot synthesis pipeline provides the data diversity that alignment alone cannot supply.

Embodied intelligence is still at an early stage. Contact-rich, long-horizon real-world tasks, failure recovery, continual learning, and complex human-robot-environment interactions remain challenging. Yet Qwen-RobotManip points to a clear path forward:

Alignment unlocks scale, and scale unlocks generalization.

@article{qwenrobotmanip, title={Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models}, author={Qwen Team}, year={2026}}

来源:Qwen:Blog Retrieval(API) · qwen.ai