跳到正文
北京时间
原文
Ollama:Blog·· 7 天前精选AI 评分69

Ollama 支持基于 Jev API 的决策模型,新增 nimble 等三款模型

Ollama now supports Jev-style decision models

AI 导读

Ollama 0.35 通过新 /v1/systemone 端点支持基于 TypeSafe Jev API 的决策模型,可在本地一次请求回答多个命名问题,适合工单分诊、模型路由和内容审核等快速决策任务。

推荐理由

官方宣布本地运行决策模型,给出延迟数据和 API 用法,可帮助读者评估本地快速决策场景的可行性。

正文 · AI 翻译

2026年9月29日

Ollama 现已支持决策模型,基于 TypeSafe 的 Jev API,实现快速、类型化的决策:

  • 无额外费用
  • 本地运行时延迟更低
  • 今天通过 Ollama 可用的三个新决策模型

这个新 API 自 Ollama 0.35 起可用,通过新的 /v1/systemone 端点使用。以 state 形式发送文本并附带一组命名问题,运行在你机器上的模型会在一次请求中回答所有问题。这非常适合需要快速决策的任务,例如工单分类、模型路由以及内容或安全审核。

近乎即时的决策

Ollama 上的决策模型速度很快,因为请求无需通过网络传输。在下方 Pac-Man 示例中,Nimble 9B 在 M5 Max 上本地运行时,平均每次决策耗时 91 毫秒。这足以快速做出决策,例如玩游戏或实时处理内容:

第 1 步

Nimble 9B 在 MacBook Pro M5 Max 上运行,以实时速度回放。

可用模型

三个新的决策模型可通过 Ollama 运行:

  • nimble:由 Bespoke Labs 开发的开源 9B 参数决策模型
  • tev1:来自 Together AI 的实验性 4B 决策模型
  • tev1:0.8b:来自 Together AI 的实验性 0.8B 决策模型

更多决策模型即将推出,包括由 Ollama 云提供的模型。

Ollama 上的决策模型

Bespoke Labs 公开基准 · 准确率,越高越好

覆盖 13 个带人工标签的公开数据集的平均准确率,涵盖 3,880 项决策。Nimble 和 Tev1 在 Ollama 上进行了评估;Jev 1.13 来自 Bespoke Labs 在相同决策上发布的运行结果。参见 Ollama 评估结果和基准测试套件。

开始使用

要开始使用,请先下载或升级到最新版本的 Ollama。接下来,下载一个决策模型,例如 nimble:

ollama pull nimble

你可以通过 curl 或 TypeSafe 官方 Python SDK 发起请求。

请求

curl http://localhost:11434/v1/systemone -d '{
  "model": "nimble",
  "state": {
    "ticket": "I was charged twice. Please refund the extra payment."
  },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this ticket?",
      "criteria": {
        "billing": "Payments and refunds",
        "technical": "Bugs and integrations",
        "other": "None of the above"
      }
    },
    "refund": {
      "type": "noul",
      "instructions": "Does the customer explicitly ask for a refund?"
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgent is this ticket?",
      "criteria": ["Routine", "Soon", "Urgent"]
    }
  }
}'

响应

{
  "model": "nimble",
  "answers": {
    "team": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
      "confidence": 0.922
    },
    "refund": {"type": "noul", "noul": 0.997},
    "urgency": {
      "type": "score",
      "score": 0.815,
      "legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
      "probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
      "confidence": 0.046
    }
  },
  "usage": {"input_tokens": 841, "output_tokens": 4}
}

下一步

这是为 Ollama 添加决策模型支持的众多版本中的第一个。未来更新将包括:

  • 由 MLX 驱动的 Apple Silicon 上更快的性能
  • 更多专注于不同类型决策的模型

来源:Ollama:Blog · ollama.com