跳到正文
北京时间
原文
Tomer Tunguz 博客(VC 分析)· Tomasz Tunguz·· 2026-07-08精选AI 评分57

AI预检检查:智能体工作记忆架构

The AI Preflight Check

AI 导读

一种为AI智能体设计的预检工作记忆架构:查询到来时,系统从磁盘上约90个索引化的技能库中检索最相关技能,仅加载到上下文窗口。本地开源模型Ornith 35B(350亿参数,通过Ollama在Apple Silicon上运行)执行任务,约80%常规任务由本地模型完成,困难任务路由至前沿模型。看门狗记录每次预检决策和技能调用,夜间通过异步推理处理全天轨迹,自动决定哪些技能需新增或固化(如日历排期转为确定性Rust代码),实现自我改进循环。昨天,看门狗首次未提出任何改进建议,系统或接近性能平台期。

推荐理由

Tunguz 把代理的记忆问题拆成预检+看门狗,不是大模型调参,而是软件架构层的优化,做 agent 的开发者可以直接偷师。

正文 · 原文

In short : A working memory architecture for AI agents : preflight retrieves the right skill from a long-term library, a local model executes, & a watchdog reads the trail overnight to update the library.

I still remember when my agent would forget what I said mid-sentence.

Context size is not the ceiling. Memory architecture is.

Diagram of the memory system : a user request flows through a preflight check that loads the right skill from a long-term skills library into the context window, executed by the local Ornith 35B model, with a watchdog reading the trail at the end.

I have been experimenting with a memory architecture that runs preflight instructions. A pilot plans the route before takeoff. My agent does the same.

A query lands. “Summarize the Q3 board deck.” 200,000 raw tokens of emails, PDFs, & chats sit behind that sentence.

Preflight is retrieval. The agent inspects its skills library1, picks the ones relevant to the task, & loads only those into the context window. Skills are consolidated memory ; the preflight step is how the agent picks the right one.

The local Ornith 35B model2 then executes on that loaded context. Hard tasks route out to the frontier ; routine tasks remain on the local model, which happens about 80% of the time.

Editorial illustration of an airline pilot in the cockpit running a preflight checklist, with a planned flight route curving through the windshield.

The watchdog monitors which skills are loaded, which decisions are made, & the success rate. Every preflight decision is logged. Every skill invocation is a named, versioned artifact.

Overnight, asynchronous inference3 processes the day’s trail. It decides which new skills should be developed, & which parts of existing skills should become deterministic code. Calendar scheduling is a good example : an LLM should not be comparing free & busy slots ; Rust is much better at that. The system rewrites its skills library & restarts itself in a self-improving loop.

Yesterday was the first day the watchdog did not suggest any improvements. I doubt it will continue. But it hints at something : at some level of improvement, the system reaches a plateau. Only genuinely new exceptions need human help.


  1. The skills library is a set of workflow files (~90 at present) indexed on-disk & retrieved by intent match. Skills are workflows written once, versioned, & handed to the model as tool schemas. See Skill Distillation for how the library was built. ↩︎

  2. Ornith 35B is a locally-hosted open-weight model in the 35-billion-parameter class, run on Apple Silicon via Ollama. It handles routine agent work — classification, drafting, tool selection, structured extraction — & routes the hard remainder to the frontier. ↩︎

  3. See Full Sail on Asynchronous Inference for the queue architecture that makes overnight, hours-long agent runs tractable. ↩︎

来源:Tomer Tunguz 博客(VC 分析) · tomtunguz.com