Google Cloud API Gateway 推出统一模型路由功能,支持 Gemini、Claude 与 OpenAI OSS-GPT
A unified API for AI model routing
Google Cloud API Gateway 新增模型路由功能(Public Preview),开发者可在 OpenAPI 3.x 规范中配置虚拟模型名到后端目标的映射,无需硬编码端点或管理开源代理。网关作为无服务器入口层,接受标准 OpenAI 兼容请求,自动将负载转码为目标模型的原生格式并动态路由流量。
将模型路由规则直接内嵌到 OpenAPI 规范中,减少了多模型调度时的入口适配工作,也无需自行维护代理层。
Model routing with Google Cloud API Gateway
Mak Ahmad
Sanjay Pujare

When building AI applications, developers need the freedom to route traffic to the best model for the job without hardcoding endpoints or managing open-source proxies. Google Cloud API Gateway now offers model routing in Public Preview to solve this. It provides a lightweight, serverless ingress layer that accepts OpenAI-compatible requests and dynamically routes them to Gemini, Claude, or OpenAI OSS-GPT. This AI gateway pattern is commonly referred to as an LLM gateway or centralized LLM endpoint.
API Gateway can be used standalone for simple rules-based routing, rate limiting and token tracking, or paired seamlessly with Google Cloud’s broader AI gateway spectrum and Gemini Enterprise Agent Platform. For example, you can route your agent's egress through Agent Gateway for strict security governance, and then pass the request to API Gateway to handle dynamic routing to Google-hosted LLMs. Here is a step-by-step guide on how to configure your routing logic.
Routing your traffic
This gives you a single, stable endpoint for all your LLM traffic, so you can add or swap backend models centrally without changing client code. And because applications authenticate to the Gateway, not to the model providers, client auth stays separate from backend LLM auth, letting you rotate or change backend credentials without touching your apps.
- Configure your routing rules: You can map virtual model names to specific backend targets directly in your OpenAPI 3.x specification using the new
x-google-api-managementextension block.
openapi: 3.0.4
info:
title: OpenAPI 3.x spec using Model Routing
description: Using Model Routing in an OAS 3.x spec
version: 1.0.0
x-google-api-management:
backends:
gemini-35-flashlite:
address: >-
https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT_ID/locations/global/publishers/google/models/gemini-3.5-flash-lite:generateContent
deadline: 60.0
pathTranslation: CONSTANT_ADDRESS
anthropic-claude-opus-47:
address: >-
https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT_ID/locations/global/publishers/anthropic/models/claude-opus-4-7:rawPredict
deadline: 60.0
pathTranslation: CONSTANT_ADDRESS
openai-gpt-oss-120b:
address: >-
https://aiplatform.googleapis.com/v1/projects/YOUR_PROJECT_ID/locations/global/endpoints/openapi/chat/completions
deadline: 60.0
pathTranslation: CONSTANT_ADDRESS
ai:
models:
routing:
routers:
# Router 1: route between Gemini (default) and Claude.
gemini-claude-router:
defaultModel:
backend: gemini-35-flashlite
targetModel: google/gemini-3.5-flash-lite
rules:
- model: "claude-opus-4-7"
backend: anthropic-claude-opus-47
targetModel: anthropic/claude-opus-4-7
# Router 2: route between OpenAI GPT (default) and Gemini.
openai-gemini-router:
defaultModel:
backend: openai-gpt-oss-120b
targetModel: openai/gpt-oss-120b-maas
rules:
- model: "gemini-3.5-flash-lite"
backend: gemini-35-flashlite
targetModel: google/gemini-3.5-flash-lite
servers:
- url: "https://my-gateway.example.com"
paths:
/v1/chat/gemini-claude:
post:
summary: "Endpoint:defaults to Gemini & Claude as an option."
operationId: "chatGeminiClaude"
x-google-model-router: gemini-claude-router
responses:
'200':
description: "OK"
/v1/chat/openai-gemini:
post:
summary: "Endpoint:defaults to OpenAI & Gemini as an option."
operationId: "chatOpenAIGemini"
x-google-model-router: openai-gemini-router
responses:
'200':
description: "OK"Note: All backends referenced by a single router must share the same host (for example, aiplatform.googleapis.com). Routing selects a different model and path on that shared Agent Platform host — it does not route across different hosts.
2. Deploy the Gateway: Deploy your updated API config so the Gateway is active and ready to process traffic.
3. Send standard requests: Your application simply sends a standard OpenAI POST /v1/chat/gemini-claude or POST /v1/chat/openai-gemini request. The Gateway intercepts it, transcodes the payload to the native schema of the backend, adds the required agent platform authentication token, and routes it on the fly. As an example (use appropriate values for $API_KEY and my-gateway.example.com) :
curl -X POST "https://my-gateway.example.com/v1/chat/gemini-claude" \
-H "content-type: application/json" \
-H "x-api-key: $API_KEY" \
-d '{
"model": "claude-opus-4-7",
"messages": [
{"role": "user", "content": "Introduce yourself in 5 words"}
]
}'Model routing is now available in Public Preview for API Gateway. To stop managing proxies and start unifying your AI traffic, check out our documentation to deploy your first model router today.
These model routing capabilities are part of Google Cloud's broader AI gateway spectrum: expansive API and tools management with Apigee, and end-to-end agent governance with Agent Gateway in the Gemini Enterprise Agent Platform, so you can start lightweight and grow into the rest when you need it.
来源:Google Developers Blog(RSS) · developers.googleblog.com