跳到正文
北京时间
原文
Jim Fan· @DrJimFan · X·· 2025-08-05精选
AI 导读

NVIDIA发布DreamGen引擎(GR00T Dreams),将Sora/Veo等视频生成模型用作神经物理引擎,通过微调模型、模拟并行世界、恢复伪动作、训练基础模型四步流程,为机器人生成大规模合成训练数据。人形机器人仅凭单一拾放任务即可学会倾倒、折叠等22种新行为,在新动词和陌生环境中实现零样本泛化(成功率分别达43%和28%)。相比传统图形引擎,该方法以恒定计算成本处理可变形物体、流体等复杂交互,团队计划数周内完全开源。

推荐理由

NVIDIA提出用视频生成模型为机器人“造梦”合成训练数据,实现零样本技能泛化

正文 · AI 翻译

机器人领域的世界建模极其困难,因为 (1) 对类人机器人及五指手的控制,远比游戏中上⬆️左⬅️下⬇️右➡️(Genie 3 那样)要复杂得多;(2) 物体交互的多样性远超完全自动驾驶(FSD),因为 FSD 需要*避免*发生接触。我们的 GR00T Dreams 工作是构建高保真类人机器人世界模拟器的首次尝试。它不仅用于评估,还用于大规模合成数据生成。是时候告别机器人领域的"化石燃料"(人工遥操作),拥抱清洁能源(核"扩散模型")了!

GR00T Dreams 之前有些低调,所以在今天这个欢乐的日子里让它重新焕发生机 ;)

引用Jim Fan@DrJimFan
如果机器人能在视频生成模型内部做梦会怎样?我们推出 DreamGen,这是一种全新引擎,它不依赖成群的人类操作员,而是借助以像素为单位的数字梦境来规模化机器人学习。DreamGen 生成海量的神经轨迹——即与电机动作标签配对的逼真机器人视频——并解锁了对新名词、新动词和新环境强大的泛化能力。无论你是人形机器人(GR1)、工业机械臂(Franka),还是一个可爱的小机器人(HuggingFace SO-100),DreamGen 都能让你做梦。 像 Sora 和 Veo 这样的视频生成模型本质上是神经物理引擎。通过压缩数十亿条互联网视频,它们学习到了一个包含无数可能的未来的多重宇宙——即从任意初始图像帧出发,世界可能如何展开的叠加态。DreamGen 通过一个简单的四步方案利用了这种能力: 1. 在目标机器人上微调一个 SOTA 视频模型; 2. 用多样化的语言提示词提示模型,模拟平行世界:你的机器人在新场景中会如何行动。过滤掉那些不遵循指令的坏梦(哈!); 3. 利用逆动力学或潜在动作模型恢复伪动作; 4. 在經過大规模增强的神经轨迹数据集上训练机器人基础模型。 就这么简单。只是更多的数据,以及纯粹的旧式监督学习。很简单,对吧? 令人惊叹的是它的效果能走多远。仅从一个单任务的“拾取-放置”数据集出发,我们的人形机器人学会了 22 种新行为,例如倒水、折叠、舀取、熨烫和锤击,尽管它从未见过这些动词。更棒的是,我们可以把机器人从实验室带出来,放到 NVIDIA 总部咖啡馆里,让 DreamGen 施展它的魔法。我们展示了真正的从零到一的泛化:新动词的成功率从 0% 提升到超过 43%,在未见过的环境中从 0% 提升到 28%。 与传统的图形引擎相比,DreamGen 并不在乎场景是否涉及可变形物体、流体、半透明材质、高接触交互或疯狂的光照。手工构建这些场景可不容易。对于 DreamGen,每个世界都只是通过扩散神经网络的一次前向传播。无论梦境多复杂,展开它所需的计算时间都是恒定的。 立即阅读我们的博客和论文!我们计划在未来几周内完全开源整个管道。链接在帖子中。
原文

What if robots could dream inside a video generative model? Introducing DreamGen, a new engine that scales up robot learning not with fleets of human operators, but with digital dreams in pixels. DreamGen produces massive volumes of neural trajectories - photorealistic robot videos paired with motor action labels - and unlocks strong generalization to new nouns, verbs, and environments. Whether you’re a humanoid (GR1), an industrial arm (Franka), or a cute little robot (HuggingFace SO-100), DreamGen enables you to dream. Video generation models like Sora & Veo are neural physics engines. By compressing billions of internet videos, they learn a multiverse of plausible futures, i.e. superpositions of how the world could unfold from any initial image frame. DreamGen taps into this power with a simple 4-step recipe: 1. Fine-tune a SOTA video model on your target robot; 2. Prompt the model with diverse language prompts to simulate parallel worlds: how your robot would have acted in new scenarios. Filter out the bad dreams (ha!) that don’t follow instructions; 3. Recover pseudo-actions using inverse dynamics or latent action models; 4. Train robot foundation models on the massively augmented dataset of neural trajectories. That’s it. Just more data, and plain old supervised learning. Simple, right? What’s remarkable is how far this goes. Starting with just a single-task dataset of pick-and-place, our humanoid robot learns 22 new behaviors, such as pouring, folding, scooping, ironing, and hammering, despite never seeing those verbs before. Better yet, we can take the robot out of the lab and drop it into the NVIDIA HQ Cafe, and let DreamGen work its magic. We show true zero-to-one generalization: from 0% success to over 43% for novel verbs, and 0 -> 28% in unseen environments. Compared to a traditional graphics engine, DreamGen doesn’t care if the scene involves deformable objects, fluids, translucent materials, contact-rich interactions, or crazy lighting. Good luck engineering those by hand. For DreamGen, every world is just a forward pass through a diffusion neural net. No matter how complex the dream is, it takes constant compute time to roll out. Read our blog and paper today! We plan to fully open-source the entire pipeline in the next few weeks. Links in thread:

在 X 查看被引用的帖子

来源:Jim Fan · x.com