跳到正文
北京时间
原文
Saining Xie· @sainingxie · X·· 2025-12-16精选
AI 导读

新论文:iREPA 扩散模型是其底层表征的渲染器。通过这种新设置,我们能更清楚地洞察这些表征的真正含义。Jas 开始了一场自发的探索,过去三个月我们学到了很多 ps. 这也是我们对一种新型线上"饮水机效应"的小实验,我很喜欢看到这种现象。让我们争论、讨论,然后用真正的努力将其转化为正经科学 [引用 @1jaskiratsingh]:‼️ 表征对生成很重要!但事实证明,我们对表征如何帮助生成的理解一直都是错的 ‼️ 我们之前的想法:(我们错了) ❌ 更大的视觉编码器 → 更好的表征 → 更好的生成 ❌ 更好的全局语义 → 更好的表征 → 更好的生成 结果发现: 🤯 在表征对齐方面,小 20 倍以上的视觉编码器可以达到与更大模型相似或更好的性能 🤯 线性探测准确率约 20%(全局语义的衡量指标)的视觉编码器可以胜过准确率 >80% 的编码器 🤯 即使是 SiFT 和 HoG 这类经典特征也能带来与现代大得多的视觉编码器相媲美的提升 ‼️ 🚨 介绍:什么对表征对齐重要?全局信息还是空间结构 🚨 TL;DR: ✅ 更好的全局语义信息 ≠ 更好的生成 ✅ 空间结构(而非全局语义)驱动表征的生成性能 ✅ 我们提出 iREPA:仅需 3 行代码,强调空间结构迁移,并在 REPA、REPA-E、Meanflow、JiT 等方法上持续提高收敛速度 在 @AdobeResearch 的激动人心的项目,与 @xingjian_leng、@zongze_wu、@LiangZheng_06、@rzhang88、@elishechtman 和 @sainingxie 合作 🙏 对我来说这也是一次特别有趣且独特的经历,在项目的每一步我们都在证明自己的偏见是错误的 😆 还要大力感谢 @YouJiacheng、@ShumingHu 和 @gallabytes,他们在 X 上的评论开启了这一方向的探索 🫡 论文:https://arxiv.org/abs/2512.10794 代码:https://github.com/End2End-Diffusion/iREPA 项目页面:https://end2end-diffusion.github.io/irepa 更多细节见线程:[1/n] 🧵

推荐理由

颠覆认知:小20倍视觉编码器也能驱动高质量生成,空间结构才是关键

正文 · 原文

new paper: iREPA

diffusion models are a renderer of their underlying representations. with this new setup, we can gain much clearer insight into what those representations are really about. Jas took on a spontaneous quest, and over the past three months we have learned so much

ps. this is also our little experiment in a new kind of online water cooler effect that I loved seeing. let’s argue, discuss, and then turn it into proper science with real effort

引用Jaskirat Singh@1jaskiratsingh
‼️ Representations matter for generation! But turns out our understanding of how representations help generation was wrong all along ‼️ What we thought: (we were wrong) ❌ Bigger vision encoders → better representations → better generation ❌ Better Global Semantics→ better representations → better generation Turns out: 🤯 >20x smaller vision encoders can have similar or better performance then much bigger models for representation alignment 🤯 Vision encoders with ~20% linear probing accuracy (measure of global semantics) can outperform encoders with >80% accuracy. 🤯 Even classical features like SiFT and HoG can give competitive gains similar to modern much larger vision encoders ‼️ 🚨 Introducing: What matters for representation alignment? Global Information or Spatial Structure 🚨 TL;DR: ✅ Better Global Semantic Information ≠ Better Generation ✅ Spatial Structure (not global semantics) drives the generation performance of representations ✅ We propose iREPA: just 3 lines of code which accentuate spatial structure transfer and consistently improve convergence speed across REPA, REPA-E, Meanflow, JiT etc. Exciting project at @AdobeResearch and collaboration with @xingjian_leng, @zongze_wu, @LiangZheng_06, @rzhang88, @elishechtman and @sainingxie 🙏 This was also a particularly fun and unique experience for me where we were proving our own biases wrong at each step of the project 😆 Also huge shoutout to @YouJiacheng, @ShumingHu and @gallabytes whose comments here on X initiated the exploration in this direction 🫡 Paper: https://arxiv.org/abs/2512.10794 Code: https://github.com/End2End-Diffusion/iREPA Project page: https://end2end-diffusion.github.io/irepa More details in the thread: [1/n] 🧵
在 X 查看被引用的帖子

来源:Saining Xie · x.com