Artificial Analysis 发布 Speech to Speech Index,OpenAI 的 GPT-Live-1 以 81.5 分(Astra 后端,medium 推理强度)排名第一,超过 Grok Voice Think Fast 2.0 High 的 81.3;Sol 后端配置得 80.1 排第三。
原文给出语音到语音模型的完整分数、速度和成本对比,读者可据此在不同后端配置间做选型判断。
OpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its delegated backend model, ahead of Grok Voice Think Fast 2.0
GPT-Live-1 is @OpenAI's new full duplex Speech to Speech model that can delegate reasoning and tool use to a backend text model while continuing the conversation. Developers stream audio in and receive speech back through the API, with the backend text model configured separately. We evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort.
Key takeaways:
➤ Speech to Speech Index: GPT-Live-1 (Astra, medium) achieves 81.5, ranking #1, while GPT-Live-1 (Sol, low) scores 80.1, ranking #3. Grok Voice Think Fast 2.0 High sits between them at 81.3
➤ Speech Agent Arena: GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% Task Success Rate, while GPT-Live-1 (Astra, medium) ranks #4 at 1,048 Elo with 87.4% task success. Gemini 3.1 Flash Live Minimal leads preference at 1,096 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.6%
➤ Tau Voice: GPT-Live-1 (Astra, medium) and GPT-Live-1 (Sol, low) take the top two spots on our agentic-performance benchmark at 67.9% and 59.3%, respectively, ahead of Grok Voice Think Fast 2.0 High at 56.5%.
➤ Big Bench Audio: GPT-Live-1 (Astra, medium) scores 90.1% and GPT-Live-1 (Sol, low) scores 89.0% on audio reasoning, behind Grok Voice Think Fast 2.0 High at 97.2% and Qwen Audio 3.0 Realtime Plus at 99.2%
➤ Speed: Average Time to First Audio on Big Bench Audio is 1.34 seconds for GPT-Live-1 (Astra, medium) and 1.24 seconds for GPT-Live-1 (Sol, low), compared with 0.70 seconds for Grok Voice Think Fast 2.0 High
➤ Cost: GPT-Live-1 (Astra, medium) costs $5.83 per hour of input audio, compared with $4.47 for GPT-Live-1 (Sol, low), including delegated backend model usage, versus $4.80 for Grok Voice Think Fast 2.0 High on our fixed Big Bench Audio pricing subset
See below for more detail ⬇️
来源:Artificial Analysis · x.com