跳到正文
北京时间
原文
Artificial Analysis· @ArtificialAnlys · X·· 2026-07-18精选AI 评分76
AI 导读

过去八天内,Grok 4.5、GPT-5.6、Muse Spark 1.1 与 Kimi K3 四款前沿模型相继发布,使 Artificial Analysis Intelligence Index 得分超 50 的实验室从 6 月初的 2 家增至 6 家。

推荐理由

八天四个前沿模型,Frontier 从两家实验室变成六家,前三名分差仅三分,而且价格正在暴跌,模型选型的逻辑从「谁第一」变成了「够用且便宜」——对产品经理来说是真正的转折点。

正文 · 原文

The frontier has opened up: four frontier launches in eight days - Grok 4.5, GPT-5.6, Muse Spark 1.1, and yesterday Kimi K3; six labs now have a model scoring >50 on the Artificial Analysis Intelligence Index, up from two in early June

@SpaceXAI's Grok 4.5 (high, Artificial Analysis Intelligence Index Score: 54) landed July 8, @OpenAI's GPT-5.6 Sol, Terra, and Luna (max: 59, 55, 51) and @AIatMeta's Muse Spark 1.1 (xhigh, 51) followed a day later, and @Kimi_Moonshot's Kimi K3 launched yesterday at 57 - third overall, ahead of Claude Opus 4.8 (max, 56).

The top three models on the Index now come from three different labs and span just three points. Four of the ten highest-scoring models launched since July 8, and six of ten since early June.

The one thing that did not move is #1: Claude Fable 5 (max, 60) has held the top spot since June 9, but its lead has narrowed from four points to one, and the price of the intelligence beneath it collapsed.

Congratulations to @elonmusk, @sama, @finkd, and the teams at all four labs on a remarkable eight days.

Key Takeaways:

➤ The frontier went from two labs to six in six weeks. Until June, only Anthropic and OpenAI had fielded a model at 51 or above. GLM-5.2 (max) brought Z AI in mid-June; last week added SpaceXAI and Meta; today Kimi K3 makes Moonshot AI the sixth as they enter at 57

➤ Kimi K3 debuts at #3 with agentic and knowledge work scores behind only the top two. K3 scores 1668 Elo on GDPval-AA v2, third behind Claude Fable 5 (max, 1760) and GPT-5.6 Sol (max, 1748). On AA-Briefcase, our benchmark of long-horizon knowledge work, it enters at #2 with 1547 Elo - behind only Claude Fable 5 (max, 1583) and ahead of GPT-5.6 Sol (max, 1495) - with an Analytical Quality Elo (1760) effectively tied with Fable 5 (1764). At $0.94 per Intelligence Index task on its $3/$15 pricing, it delivers comparable intelligence to Claude Opus 4.8 (max, $1.80) at roughly half the cost per task

➤ Near-frontier intelligence got 2-3x cheaper in eight days. GPT-5.6 Sol (max) delivers one point below Claude Fable 5 (max) at $1.04 per Intelligence Index task vs $2.75. Grok 4.5 (high) delivers 54 at $0.31, under a third of GPT-5.5 (xhigh, $0.99). At 51, GPT-5.6 Luna (max, $0.21) and Muse Spark 1.1 (xhigh, $0.26) undercut GLM-5.2 (max, $0.32), the cheapest at that level a week earlier

来源:Artificial Analysis · x.com