Ollama:Blog·· 7 天前精选AI 评分69
Ollama 支持基于 Jev API 的决策模型,新增 nimble 等三款模型
Ollama now supports Jev-style decision models
AI 导读
Ollama 0.35 通过新 /v1/systemone 端点支持基于 TypeSafe Jev API 的决策模型,可在本地一次请求回答多个命名问题,适合工单分诊、模型路由和内容审核等快速决策任务。
推荐理由
官方宣布本地运行决策模型,给出延迟数据和 API 用法,可帮助读者评估本地快速决策场景的可行性。
正文 · AI 翻译
2026年9月29日
Ollama 现已支持决策模型,基于 TypeSafe 的 Jev API,实现快速、类型化的决策:
- 无额外费用
- 本地运行时延迟更低
- 今天通过 Ollama 可用的三个新决策模型
这个新 API 自 Ollama 0.35 起可用,通过新的 /v1/systemone 端点使用。以 state 形式发送文本并附带一组命名问题,运行在你机器上的模型会在一次请求中回答所有问题。这非常适合需要快速决策的任务,例如工单分类、模型路由以及内容或安全审核。
近乎即时的决策
Ollama 上的决策模型速度很快,因为请求无需通过网络传输。在下方 Pac-Man 示例中,Nimble 9B 在 M5 Max 上本地运行时,平均每次决策耗时 91 毫秒。这足以快速做出决策,例如玩游戏或实时处理内容:
第 1 步
可用模型
三个新的决策模型可通过 Ollama 运行:
nimble:由 Bespoke Labs 开发的开源 9B 参数决策模型tev1:来自 Together AI 的实验性 4B 决策模型tev1:0.8b:来自 Together AI 的实验性 0.8B 决策模型
更多决策模型即将推出,包括由 Ollama 云提供的模型。
Ollama 上的决策模型
Bespoke Labs 公开基准 · 准确率,越高越好
开始使用
要开始使用,请先下载或升级到最新版本的 Ollama。接下来,下载一个决策模型,例如 nimble:
ollama pull nimble
你可以通过 curl 或 TypeSafe 官方 Python SDK 发起请求。
请求
curl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {
"ticket": "I was charged twice. Please refund the extra payment."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}'响应
{
"model": "nimble",
"answers": {
"team": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
"confidence": 0.922
},
"refund": {"type": "noul", "noul": 0.997},
"urgency": {
"type": "score",
"score": 0.815,
"legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
"probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
"confidence": 0.046
}
},
"usage": {"input_tokens": 841, "output_tokens": 4}
}下一步
这是为 Ollama 添加决策模型支持的众多版本中的第一个。未来更新将包括:
- 由 MLX 驱动的 Apple Silicon 上更快的性能
- 更多专注于不同类型决策的模型
来源:Ollama:Blog · ollama.com