PromptArmor:供应商接入 Fireworks 等推理商后往往跟进 Kimi、DeepSeek、GLM 等国际模型
Vendors adding Fireworks? Kimi, DeepSeek and GLM are on the way
PromptArmor 分析称,供应商因成本优势和基准竞争力正转向国际与开源权重模型,DeepSeek V4 Flash 输入价 $0.14/M tokens,远低于 GPT-5.6 Sol 的 $5。
作者基于自家监测数据提出可迁移的预判方法,供应商接入 Fireworks 等推理商后往往跟进国际模型,可帮助读者预判供应链变化。
![]()
Your vendors
Glean
Harvey
Zendesk
Which of your vendors are adopting new inference hosts and international models?
Vendors are turning to international AI models
Vendors with AI features are increasingly turning to international and open-weight models due to the cost difference compared to models from leading U.S. labs, and because these international models are becoming increasingly competitive on benchmarks.
Recently, Kimi K3 and DeepSeek V4 Flash 0731 have made the news, offering competitive performance at a steep discount compared to models like Anthropic's Fable and OpenAI's GPT Sol.
International vs U.S. model token costs
International model
U.S. model
Input, then output
![]()
DeepSeek
V4 Flash
in $0.14
![]()
DeepSeek
V4 Pro
in $0.435
![]()
Z.ai
GLM 5.2
in $1.40
![]()
Mistral
Medium 3.5
in $1.50
![]()
OpenAI
GPT-5.6 Terra
in $2.00
![]()
Moonshot
Kimi K3
in $3.00
![]()
Anthropic
Sonnet 5
in $3.00
![]()
Anthropic
Opus 5
in $5.00
![]()
OpenAI
GPT-5.6 Sol
in $5.00
![]()
Anthropic
Fable 5
in $10.00
International
U.S.
![]()
DeepSeek V4 Flash
In
$0.14
Out
$0.28
![]()
DeepSeek V4 Pro
In
$0.435
Out
$0.87
![]()
Z.ai GLM 5.2
In
$1.40
Out
$4.40
![]()
Mistral Medium 3.5
In
$1.50
Out
$7.50
![]()
OpenAI GPT-5.6 Terra
In
$2.00
Out
$12.00
![]()
Moonshot Kimi K3
In
$3.00
Out
$15.00
![]()
Anthropic Sonnet 5
In
$3.00
Out
$15.00
![]()
Anthropic Opus 5
In
$5.00
Out
$25.00
![]()
OpenAI GPT-5.6 Sol
In
$5.00
Out
$30.00
![]()
Anthropic Fable 5
In
$10.00
Out
$50.00
USD per million tokens. Provider list prices, read August 9, 2026. Sources: DeepSeek, Z.ai, Mistral, Moonshot, OpenAI, Anthropic.
Furthermore, open-weight models do not impose API-level restrictions on security and biology work, unlike leading U.S. model providers.
More on How Hugging Face Turned to Zhipu due to Claude's Guardrails

Hugging Face's incident responders were blocked from analyzing their own intrusion because the attack commands in the logs tripped Anthropic's safety guardrails. They moved to Zhipu AI's GLM 5.2 to complete their investigation.
While organizations are not all aligned on the potential for risks such as model backdoors, most organizations are in agreement that data processing for AI inference must remain within the U.S.
Because of this, vendors that want to add international or open-weight models turn to inference providers whose job it is to run those models within U.S. data residency.
Inference provider changes predict model changes
Across the vendors we monitor for our customers, we see AI subprocessor changes frequently. From new features that come with new models to updating features with the latest models, subprocessor and model changes are common occurrences. Throughout these changes, there have been some notable trends.
One sequence we have seen play out again and again is that shortly after a vendor adds a US-based inference provider like Fireworks, Baseten, Together, Groq, or Cerebras, they begin rolling out open-weight or international models.
Below are samples of vendors we have seen take this path, and vendors who have recently added an inference provider but are yet to announce their more-than-likely function: supporting the adoption of international or open models.
Recent examples from vendors
Vendor
Subprocessor added
Date
Model added
Date
International models following new inference subprocessor
Glean
Fireworks
Jun 16, 2026
![]()
Z.ai GLM 5.2
Jul 16, 2026
Harvey
Baseten
Fireworks
Jul 3, 2026
![]()
Mistral Medium 3.5
Aug 5, 2026
Vercel
Cerebras
Baseten
Jan 29, 2026
![]()
Z.ai GLM 5.1
Apr 8, 2026
GitHub Copilot
Fireworks
Aug 29, 2025
![]()
Moonshot Kimi K3
Aug 6, 2026
New inference subprocessor, new models coming soon
Sierra
Baseten
Fireworks
Together
Modal
Jun 3, 2026
Coming soon
Zendesk
Groq
Fireworks
Baseten
Jul 17, 2026
Groq, Dec 2025; Fireworks, Mar 2026; Baseten added and Fireworks broadened to hosting and inference, Jul 17, 2026
Coming soon
Ada
Groq
Baseten
Jun 2026
Groq, Jan 2026; Baseten, Jun 2026
Coming soon
Deepgram
Baseten
May 2026
Coming soon
International models following new inference subprocessor
Glean
Subprocessor added
Jun 16, 2026
Fireworks
Model added
Jul 16, 2026
![]()
Z.ai GLM 5.2
Harvey
Subprocessor added
Jul 3, 2026
Baseten
Fireworks
Model added
Aug 5, 2026
![]()
Mistral Medium 3.5
Vercel
Subprocessor added
Jan 29, 2026
Cerebras
Baseten
Model added
Apr 8, 2026
![]()
Z.ai GLM 5.1
GitHub Copilot
Subprocessor added
Aug 29, 2025
Fireworks
Model added
Aug 6, 2026
![]()
Moonshot Kimi K3
New inference subprocessor, new models coming soon
Sierra
Subprocessor added
Jun 3, 2026
Baseten
Fireworks
Together
Modal
Model added
Coming soon
Zendesk
Subprocessor added
Jul 17, 2026
Groq, Dec 2025; Fireworks, Mar 2026; Baseten added and Fireworks broadened to hosting and inference, Jul 17, 2026
Groq
Fireworks
Baseten
Model added
Coming soon
Ada
Subprocessor added
Jun 2026
Groq, Jan 2026; Baseten, Jun 2026
Groq
Baseten
Model added
Coming soon
Deepgram
Subprocessor added
May 2026
Baseten
Model added
Coming soon
More predictive alerts for AI in vendors
Beyond the identification of international and open-weight model use, we have observed a number of trends that can be extrapolated to inform decisions about AI in vendors. Here are a few:
Cyber-capable model adoption. Anthropic added a “Covered Models” section to its Service Specific Terms on June 8, 2026, one day before launching Claude Fable 5 as a covered model on June 9, with supplemental terms that permit human safety review and explicitly supersede zero data retention commitments for their models that have advanced cyber capabilities. The upstream data processing requirements became a signal for the adoption of these models. For example, Glean introduced a Limited Retention Addendum on June 10, two days after Anthropic’s terms change, and added Claude Fable 5 within a month.
AI feature expansion. Across industries, adoption of new tools and AI capabilities usually occurs in groups. For example, Skills first spread across coding agents, and then shortly after they spread across knowledge work. Once one major player in an industry adopts a trending AI capability, competitors are quick to adopt - so by tracking the first entrant in an industry, we can predict when other industry players will adopt. This is playing out right now with Legal AI; Legora introduced Skills on June 3, 2026; since then, Harvey and other competitors have begun rolling out Skills. This pattern has played out multiple times now, including with Memory, Web Search, MCP, and recently, 'Coworkers' and always-on-agents.
Agentic platform access via nth party apps. As applications are increasingly accessed via AI connectors, vendors have begun to shift their posture to reflect the utilization of their services via third-party AI services, like Claude and ChatGPT. Vendors have begun to release terms changes that indicate when and how users are liable for access to their services via third-party AI, such as MCP servers and connectors. In addition to legal changes, a common tendency is to have a staggered rollout of ChatGPT and Claude connectors: if we see a vendor become accessible via ChatGPT or Claude, we can infer that it will soon be accessible through the other.
Vendors are Deflecting Liability for Anything 'AI'
Agentic scope expansion. Vendors love to describe their features as 'agents' and 'agentic', but it is not uncommon to find 'agents' with limited practical capabilities beyond a normal chatbot. In the process of reviewing these agents, we've come across a pattern: from the first time an 'agentic' system surfaces in marketing, it takes about six months to go from a 'chatbot with a specialized prompt' to true agentic, autonomous, or always-on capabilities (MCP, web browsing, recurring workflows, code execution, etc.). Monitoring 'agents' for this transition has become a priority for many organizations that have approved these 'agents' on the basis of the limited toolset they originally presented.
Connector and integration releases. Recently, we released research that found that Claude and ChatGPT connectors change on average every 9 minutes. But the big vendors are not the only ones with tumultuous connector ecosystems. Across vendors of varying sizes with AI features, we’ve seen that the overhead to create one connector vastly lowers the prerequisites for future connectors. Shortly after we detect a vendor adding its first connector, they often enter a period of rapid growth in which dozens or hundreds of connector options are added, before the increase tapers back down and lands at a steady rate.
Claude and ChatGPT Connectors Change Every 9 Minutes
We monitor tens of thousands of vendors across industries, which has enabled us to convert insights across tools into predictive alerts. This helps our customers stay ahead of the curve as their vendors expand and change their AI posture.
来源:PromptArmor:Threat Intelligence · promptarmor.com