跳到正文
北京时间
原文
François Chollet· @fchollet · X·· 26 天前精选AI 评分81
AI 导读

François Chollet 发文称 GPT-6 Astra 在交互式推理任务上带来阶跃式能力提升,使用标准 harness 在 ARC-AGI-3 上得 66%,配合持续对话 harness 和自定义 compaction 接近 100%,每局成本约 $360。

推荐理由

ARC Prize 作者基于自家标准 harness 的实测数据评估 GPT-6 Astra,读者可对比 66% 与近 100% 两种设置看模型与 harness 能力的边界变化。

正文 · 原文

GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.

Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.

We see Astra as a major breakthrough in model intelligence.

Read our post on Astra and what these results mean: https://arcprize.org/blog/astra

来源:François Chollet · x.com