ARC Prize 实测 o3 与 o4-mini 在 ARC-AGI 上的表现
Analyzing o3 and o4-mini with ARC-AGI
ARC Prize Foundation 首次公布 OpenAI o3 与 o4-mini 在 ARC-AGI 上的公开测试结果。o3-low 在 ARC-AGI-1 Semi Private Eval 得分 41%,o3-medium 达 53%,两者在 ARC-AGI-2 上均未超过 3%;o4-mini-low 在 ARC-AGI-1 得 21%。
ARC Prize 公布了自家基准上 o3 与 o4-mini 的完整实测数据和高推理档的失真问题,读者可借此校准对推理档位选择的预期。
Analyzing o3 and o4-mini with ARC-AGI ARC Prize Foundation is a nonprofit committed to serving as the North Star for AGI by building open reasoning benchmarks that highlight the gap between what’s easy for humans and hard for AI. The ARC‑AGI benchmark family is our primary tool to do this. Every major model we evaluate adds new datapoints to the community’s understanding of where the frontier stands and how fast it is moving. In this post we share the first public look at how OpenAI’s newest o‑series models, o3 and o4‑mini, perform on ARC‑AGI. Our testing shows: o3 performs well on ARC-AGI-1 - o3-low scored 41% on the ARC-AGI-1 Semi Private Eval set, and the o3-medium reached 53%. Neither surpassed 3% on ARC‑AGI‑2. o4-mini shows promise - o4-mini-low scored 21% on ARC-AGI-1 Semi Private Eval, and o4-mini-mediumhigh` reasoning setting did not return enough task completions to support reliable scoring. In most cases, the models failed to respond or timed out, leaving us with incomplete data that falls short of the bar required for leaderboard reporting. What did return introduces another complication: the first tasks to complete showed higher accuracy than those that came back later, suggesting a non-random subset to analyze. In addition to this, we found that the tasks that didn’t return on high compute tended to be less likely to be solved by lower compute models. Reporting these results would likely inflate the model’s true capabilities and misrepresent performance. However, in the spirit of transparency, when using “high” reasoning we observed: o3-high ARC-AGI-1 Semi Private Eval: Responded to 37 out of 100 tasks, 82% accuracy. ARC-AGI-2 Semi Private Eval: Responded to 15 out of 120 tasks, 6% accuracy. o4-mini-high ARC-AGI-1 Semi Private Eval: Responded to 29 out of 100 tasks, 89% accuracy. ARC-AGI-2 Semi Private Eval: Responded to 11 out of 120 tasks, 18% accuracy. To reiterate, the small number of returned tasks and the skewed solve rates make these results unrepresentative and should not be reported on. At best, they reflect an upper bound on performance under high-effort settings. We expect broader testing to bring these scores down as more challenging tasks are attempted. o3-medium is currently the strongest publicly available model we've tested. o4-mini isn't the most accurate, but it's the most cost-efficient. As always, all responses for public tasks are available on Hugging Face, and you can reproduce these runs using our Model Baseline testing harness. To view these scores in context of other models, see the ARC Prize Leaderboard. While typical single chain-of-thought (CoT) systems cluster around a ARC-AGI-1 performance ceiling of 30%, o3-medium achieves double that performance. This significant improvement isn't easily explained by simply scaling up earlier base models or standard CoT approaches. One possibility is that o3 employs an enhanced processing model or advanced sampling and optimization techniques that manage to boost accuracy without sacrificing inference speed. However, without explicit architectural insights, this remains speculative. Key Observations To try and understand why o3-high failed to respond to certain tasks, we analyzed its token usage, runtime, and performance across other models and ARC-AGI evaluations. We observed 3 key takeaways: Early responses showed higher accuracy Higher reasoning can be inefficient Minimal variance in tokens per second Early responses showed higher accuracy We noticed that tasks which the model returned sooner had higher accuracy. Those that took longer, either in duration or token usage, were more likely to fail. This signals that the model comes to a conclusion or has higher confidence for easier tasks earlier in the CoT process. As an aside, this pattern also hints that task difficulty might be inferred from a model's behavior beyond a simple correct/incorrect label. Below we show the success and token counts over time for tasks responded. Top: Accuracy declines as response time increases. Bottom: Incorrect answers tend to consume more tokens. o3-high histogram displays fewer data points due to lack of responses. Higher reasoning can be inefficient When comparing o3-medium and o3-high on the same tasks, we found that o3-high consistently used more tokens to arrive at the same answers. While this isn’t surprising, it highlights a key tradeoff: On easy tasks, o3-high often offers no accuracy gain but incurs a higher cost. If you’re cost-sensitive, evaluate if you really need to use high reasoning, medium may be the better default. However, if maximizing accuracy is critical and cost is less of a concern, high reasoning still has its place. Blue points above the dotted line represent tasks that both medium and high solved, but high used more tokens to solve the same problem—signaling higher cost for equivalent output. Minimal variance in tokens per second Next, we looked at tokens-per-second for each task across o-series models. We found that o3-mini-low and o4-mini-low had higher throughput (tok/s) than their medium and high counterparts. This indicates a likely algorithmic differences in mini models, though its exact cause remains unclear. Mapping tokens per second for each task across models. Notes: o1-mini does not have reasoning effort, o1-pro did not return meaningful data to report. Help Evaluate Frontier Systems OpenAI’s newest o‑series releases push the boundaries of reasoning models, they keep the frontier moving and give the community visibility into what today’s models can (and still can’t) do. ARC‑AGI exists to serve as the guidepost that shows how far we’ve come. As these systems grow more powerful, efficiency (how fast, at what cost, and using how few tokens a model solves problems) becomes the key differentiator. If you’re excited to contribute frontier model analysis or help fund transparent, public benchmarks, we’d love to talk. Reach us at team@arcprize.com or consider supporting the ARC Prize Foundation today. Thank you to Henry Pinkard for leading data analysis, Mike Knoop for a review of an early draft and OpenAI for credits to perform additional testing.
来源:ARC Prize:官方博客 · arcprize.org