跳到正文
北京时间
原文
Saining Xie· @sainingxie · X·· 2025-11-27精选
AI 导读

Meta研究人员透露,Facebook自2020年起使用TPU训练AI,由Kaiming He领导开发TF和JAX代码库,MAE、DiT等模型完全基于TPU构建。因内部采用有限,Meta于2023年取消GCP协议。推文指出,Google、Anthropic等实验室长期使用TPU训练大模型,Nvidia的CUDA护城河并非不可逾越,OpenAI亦投资Triton寻求替代。TPU与GPU的效率差异并非关键,系统工程人才才是决定性因素。

推荐理由

何恺明团队2020年起用TPU训练MAE/DiT,Nvidia护城河比想象更浅

正文 · AI 翻译

大多数人都不知道这一点,我们早在 2020 年就在 *Facebook* 使用了 TPU。

Kaiming 领导了 TF 和 JAX 代码库的初步开发,而像 MAE、MoCo v3、ConvNeXt v2 和 DiT 这样的研究项目则*完全*在 TPU 上开发。

因为我们是 FAIR 中唯一使用它们的团队,Meta 在 2023 年初取消了 GCP 合作。

TPU 也为我们在 NYU 的大部分大规模工作提供了支持,包括 SiT、Cambrian1/S 以及最近的 RAE、FreeFlow。

学习这套基础设施需要经历大量痛苦(这不是他们当初所期望的,但我的学生们现在基本都成了 TPU/JAX/XLA 专家),然而一旦掌握了,其性能和稳定性都极为出色。

对 Google 发展 TPU 和 JAX 生态系统并推动其商业化落地感到非常乐观。

引用Clive Chan@itsclivetime
我不断看到关于 TPU 的东西,有什么实质性的新进展吗? 没有证据表明 Google 曾在非 TPU 硬件上训练过 Gemini,这可以追溯到多年前 GPT 之前的模型,比如 BERT。TPU 比 Nvidia 自己的张量核心还要早。 Anthropic(以及 Character、SSI 和 Midjourney……)长期以来一直在使用 TPU,如果 Meta *没有* 在考虑它们,我会感到惊讶。 Nvidia 的护城河对大型实验室来说从来都不深——看看 OpenAI 决定它可以做得比 CUDA 更好,转而投资 Triton,在基准测试中经常胜过 CUDNN。这一切都没有什么神奇或结构性的东西,只是优秀的工程师在做优秀的工作。 TPU 并不比 GPU 高效多少,而微小的性能/功耗差异与 Meta 是否拥有正确的内核/系统工程人才来做到这一点相比,简直微不足道。Nvidia 和 Google 的护城河都很小,我们仍然处于个别优秀工程师就能扭转整个平衡的阶段。见下面的旧帖。 这难道……没有被定价进去吗?这些都是超级古老的公开信息。
原文

I keep seeing stuff about TPU, has anything materially new happened? There’s no evidence Google has ever trained a Gemini on non-TPU hardware, going years back to pre-GPT models like BERT. TPUs predate Nvidia’s own tensor cores. Anthropic (and Character, and SSI, and Midjourney, …) have long used TPUs, I’d be surprised if Meta *weren’t* looking at them. Nvidia’s moat has never been deep for the big labs - see OpenAI deciding it could do better than CUDA and investing in Triton instead, regularly edging out CUDNN on benchmarks. There is nothing magical or structural about any of this, just good engineers doing good work. TPUs are not all that more efficient than GPUs, and small perf/W differences are dwarfed by whether Meta has the right kernels / systems engineering talent to pull it off. Both Nvidia’s and Google’s moats are small and we still are at the point where individual good engineers can flip the entire balance. See old thread below. Was this just… not priced in? This is all super old public info.

在 X 查看被引用的帖子

来源:Saining Xie · x.com