跳到正文
北京时间
原文
Gary Marcus:The Road to AI We Can Trust(RSS)· Gary Marcus·· 25 天前精选AI 评分74

Gary Marcus 评 GPT-6 Astra:进步明显但鲁棒性与可监控性存疑

Hot take on GPT-6 Astra

AI 导读

Gary Marcus 发文点评 GPT-6 Astra,称多项报告显示其为真正的进步,OpenAI 产品显式创建并操纵符号世界模型,令其近十年的主张获得印证。

推荐理由

作者结合自身近十年主张神经符号世界模型的立场,指出 Astra 的关键未知在鲁棒性与可监控性,判断有具体依据。

正文 · 原文

An impressive system that can (to some unknown extent) build symbolic world models

ARC Prize@arcprize GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: Image 3 7:39 PM · Sep 3, 2026 · 293K Views * * * 51 Replies · 194 Reposts · 1.81K Likes Hot take on OpenAI GPT-6 Astra*1, with a challenge to Greg Brockman’s claims about it being AGI toward the end:

  • Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.

  • As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations.

ARC Prize@arcprize Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: "L8: hub q2 (8↓). Lengths: 14=1…" - It mapped operations to exact controls: … Image 5 7:39 PM · Sep 3, 2026 · 19K Views * * * 2 Replies · 5 Reposts · 153 Likes

  • What we don’t know is how robust that capability is. That is THE key question.

  • Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.

  • And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.

  • Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.

  • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.

  • The new system appears to be lessmonitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable, not sure why.

  • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)

This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.

来源:Gary Marcus:The Road to AI We Can Trust(RSS) · garymarcus.substack.com