Google AI 发布 Gemini 3.5 Live Translate,支持 70+ 语言、近实时延迟的语音到语音翻译。该模型直接处理原始音频流,保留说话者语调、节奏和音高。东南亚超级应用 Grab 正探索将其用于司机与乘客间的跨语言沟通,其用户每月发起超 1000 万次语音通话。开发者可通过 Gemini Live API 集成 LiveKit、Fishjam、Pipecat 或 Vision Agents 构建应用。LiveKit 已实现虚拟会议室多语言即时理解;Software Mansion 结合 MoQ 协议突破流媒体瓶颈;VisionAgents AI 展示了动态多语言切换能力。开发者可在 Google AI Studio 试用并获取 Cookbook 示例代码。
Google把Gemini 3.5的语音翻译做到了近实时、保持语调,这对出海产品是个很实际的升级,实时通话翻译终于从demo走向可用。
http://x.com/i/article/2067347903761485824
How Developers are Building with Gemini 3.5 Live Translate
70+ languages. Fluid, natural speech-to-speech. Near real-time latency.
Gemini 3.5 Live Translate is fundamentally changing how developers build global, multilingual applications. By processing raw audio streams rather than converting text turn-by-turn, it preserves speaker intonation, pacing, and pitch.
Southeast Asia's leading superapp, Grab, is exploring how to break down language barriers between drivers and travelers. With users making over 10 million voice calls a month, the model makes cross-lingual communication flow naturally.
You can build and deploy your own voice translation apps by using the Gemini Live API with LiveKit, Fishjam, Pipecat, or Vision Agents. These integrations handle the complex real-time media streaming infrastructure, so you can focus on the user experience.
Take a closer look at how you build with these tools ⬇️
LiveKit (@livekit)
By combining LiveKit Agents with Gemini 3.5 Live Translate, the LiveKit team built a virtual meeting room where everyone speaks their native language and instantly understands each other, breaking down language barriers. The model handles the multilingual inputs automatically and streams speech continuously, delivering fluid translations just a few seconds behind the speaker.
Software Mansion (@swmansion)
The Software Mansion team paired Gemini 3.5 Live Translate with the ultra-fast MoQ (Media over QUIC) protocol. This integration bypasses traditional streaming bottlenecks, setting a new standard for delivering high-quality, speech-to-speech translation to large audiences with exceptionally low latency.
VisionAgents AI (@visionagents_ai)
Real-time translation is challenging, but switching between different languages dynamically is even harder. Watch as VisionAgents AI tests the limits of the new model by throwing multiple languages at it on the fly. Using Gemini 3.5 Live Translate's auto-detection, the agent effortlessly handles the rapid context switching while maintaining natural intonation, rather than falling back on robotic, turn-based responses.
Ready to build?
You can try Gemini 3.5 Live Translate in Google AI Studio and grab starter code from the Gemini Cookbook. Read the blog for more details.
来源:Google AI Developers · x.com