Noam Brown· @polynoamial · X·· 2026-03-11精选
AI 导读
当今前沿推理模型的训练路径与 AlphaGo 高度一致:先模仿大量人类数据,再扩展推理计算(从蒙特卡洛树搜索到思维链),最后用强化学习突破模仿上限。Demis Hassabis 称,十年前 AlphaGo 的"第37步"预示 AI 可攻克真实科学难题,这些思路对构建 AGI 仍至关重要。
推荐理由
Meta 研究员揭示推理模型与 AlphaGo 的技术传承,点明 RL 超越模仿的核心路径
正文 · 原文
The recipe behind today’s frontier reasoning models is surprisingly similar to AlphaGo:
1) Imitate large amounts of human data
2) Scale inference compute to reason better (back then it was Monte Carlo Tree Search, today it's Chain of Thought)
3) Use RL to go beyond imitation
Ten years ago, AlphaGo’s legendary match in Seoul heralded the start of the modern era in AI. Its famous ‘Move 37’ signaled to us that AI techniques were ready to tackle real-world problems in areas like science - and ideas inspired by these methods are critical to building AGI https://t.co/8EibfAByaG在 X 查看被引用的帖子
来源:Noam Brown · x.com