跳到正文
北京时间
原文
Goodfire Research·· 1 天前精选AI 评分72

Goodfire 为 Kimi K3 和 GLM 5.3 训练并部署生产级网络安全监控器

Training and Deploying Production Cyber Monitors on Kimi K3

AI 导读

Goodfire Research 为 Kimi K3 和 GLM 5.3 构建基于激活探针加 LLM judge 的监控级联,并部署到生产推理栈。

推荐理由

原文给出监控级联的完整数据、成本和部署细节,读者可以据此评估实时监控开放模型智能体的可行做法。

正文 · 原文

Introduction

Cybersecurity is one of the hardest settings for monitoring AI agents, because benign and malicious behavior often look similar: there’s significant overlap between the tasks involved in auditing a codebase for vulnerabilities vs. exploiting it. Distinguishing them means tracking intent across the whole trajectory, which can stretch to millions of tokens.

Despite this difficulty, cybersecurity monitoring is critical, as demonstrated by recent incidents involving agentic coding models.

The obvious solution is to have an LLM judge read every turn, but that's too slow and expensive to run alongside the agent in real time. So in practice monitoring is asynchronous: by the time a harmful trajectory is flagged, the agent has already acted.

To address this, we built probe-based cyber monitors for Kimi K3 and GLM 5.3 and deployed them on a production inference stack. As we described in our previous post, a probe reads the model's internal activations and acts as a cheap first-line filter. It sends only suspicious exchanges to an LLM judge. Compared with running the judge on every turn, this cascade:

  • is more accurate: it Pareto-dominates the judge and matches its recall at our operating point (93% recall at 5.5% FPR).
  • costs ~50x less, at under $200 per million exchanges.
  • is much faster, maintaining baseline inference token throughput and adding no latency.
  • reduces universal jailbreaks to 0% on an independent evaluator’s static battery of red-teaming attacks.

These results make it practical to monitor agentic rollouts in real time, meaning that risky cyber actions can be blocked before any tool is called. In this post, we share details on how we built these monitors, as well as some of the inference optimizations to our monitoring stack that enable such efficient performance.

Open models don't come with a full safety stack; we think a monitoring setup like this is a practical default for anyone serving them in agentic settings. If you serve or build on open models and want to deploy monitors like these, get in touch.

Our Monitor Recipe

Data

Agentic sessions can span hundreds of thousands to millions of tokens, and harmful behavior can emerge at any point. To model production data of this nature, we simulate users with various harmful and benign goals within both cyber and non-cyber (e.g., SWE, tool-use, etc.) environments. For this purpose, we collected a portion of our data using Prime Intellect’s verifiers, Prime Sandboxes, and Tunnels, allowing us to run Kimi K3 inference on cyber offensive as well as benign tasks without network access in a secure fashion.

To specify which behaviors our monitor should flag, we use a precise policy that describes undesirable cyber activity, using concrete factors that distinguish well-intended tasks (e.g., bounded vulnerability reproduction) from ill-intended ones (e.g., vulnerability exploitation). The simulated user can introduce or escalate harmful intent at any point throughout the session, including after an initially benign request. For example, most tasks start with a minimal level of benign exploration (e.g., searching a file directory), and it’s only over a course of actions that harmful activity starts to take place. This produces a diverse set of interactions from the same base tasks.

Architecture: a monitor cascade

Ideally, we could monitor the target agent by having another equally capable LLM judge classify each turn in accordance with a misuse policy. However, as rollouts get long, using another model can result in extremely high costs and compute overhead: based on OpenRouter rates, monitoring one million turns from Kimi K3 rollouts using Opus 5 would cost on the order of $200K, or via Kimi K3 itself ~$30K.

To address this problem, we exploit the fact that in a long-running agentic session, most of the agent’s activity is benign, even if the intended goal is harmful—so many agent turns need not even be processed by the judge. The challenge lies in knowing which turns warrant being looked at by the judge. To this end, in line with our own past work and recent work from othersSee prior work for related ideas on probe-based monitoring: Building Production-Ready Probes for Gemini; Detecting High-Stakes Interactions with Activation Probes; Constitutional Classifiers++. Amongst these, the turn-level classification process has been tried in Constitutional Classifiers++, but only for short rollout tasks focused on CBRN threats; the Gemini work focused on only input monitoring and hence ignores the technical challenges noted above for full input-output monitoring; and the high-stakes detection paper focuses on smaller models., we use probes to amortize the cost of using an LLM judge as a monitor.

In particular, we design a novel probing architecture that is specifically designed for capturing behaviors represented in both a localized (i.e., localized to a specific chunk in a turn or rollout) and distributed (i.e., persistently present over an entire turn) manner, while also being efficient to run, as discussed in the section below.

The probe acts as the first-line, efficient, always-on detector in a monitor cascade: it helps read whether an input, tool call, or output is high-stakes, and if the probe readout exceeds a calibrated threshold, the rollout is escalated to the LLM judge for further evaluation. After calibration, we find our probe achieves 0% FNR and 95% TNR on a held-out set of agentic and chat interactions, ensuring that the LLM judge is only called to process harmful activity. This also allows us to reduce cost by using a faster, even if less capable, judge.

Flowchart: an incoming request goes to Kimi K3, whose internal activations feed the probes inside the monitor cascade. Below threshold, the interaction continues; above threshold, it goes to the judge, which either refuses the request or continues the interaction.
Figure 1: A high level overview of our monitoring pipeline. Our probes receive internal activations from the target agent (Kimi K3 in this case) as the inference server produces tokens. Probe decisions are flagged based on a threshold to the LLM judge, which can refuse the request or continue the interaction.

Inference Optimizations

Running a forward pass on a probe is far less expensive than an LLM judge. However, making sure it adds no latency while running inside a production inference server requires careful engineering work. Modern inference servers like SGLang and vLLM achieve high throughput by processing many requests together, reusing earlier prefix computations, and generating several candidate tokens at a time with speculative decoding. The monitor must preserve these throughput gains while keeping scores aligned with the correct tokens and sessions.

In our previous interpretability infrastructure post, we described how to collect activations for analysis.Goodfire, “Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model.” For the live monitoring described in this post, we extended SGLang to compute probe scores directly inside the serving engine.

We made several optimizations to preserve inference throughput while running our probes, including the implementation of custom kernels for our probe logic. We also maintain probe scores through the radix tree, DSpark speculator, and prefill-decode disaggregation to avoid additional recomputation. When serving multiple Kimi K3 model replicas, each worker runs the probes and maintains histories for its active requests. We extended the SMG router to preserve our monitoring configuration through request conversion and pass it to the selected worker. Monitoring events are returned with the response stream, allowing downstream consumers to handle events regardless of which replica served the request.

Table of throughput change with monitoring relative to serving without probes, median of three repeats. Output monitor / input plus output monitor: short workload +1.0% / −1.1%, medium +0.2% / −1.1%, shared prefix −0.6% / −1.4%, long −0.3% / −1.1%.
Figure 2: Difference in throughput when running our probes during inference compared to default SGLang, across four different inference workloads. We see a less than 2% decrease, which is within noise for our throughput testing. We measure throughput with probes on output tokens alone, and on both input and output tokens.

Overall, we find that running our probes during inference adds no overhead across several different workloads. Across four inference workloads, enabling the probes maintained throughput within the noise threshold of our benchmarks relative to default SGLang serving.

Bar chart of mean added latency per turn: LLM judge with reasoning 2,750 ms (bar truncated), LLM judge 160 ms, our cascade 42 ms, and our cascade without reasoning 2 ms.
Figure 3: Average time added to each agent turn by monitoring. By sending only flagged exchanges to the judge, our cascade reduces this delay by 50× compared with judging every turn and adds very little latency to default SGLang serving.

We also evaluate the added latency of our monitor cascade, comparing it with an LLM judge reviewing every turn. Without judge reasoning, our cascade reduces average added latency per turn from 160 ms to 2 ms; with reasoning, it reduces latency from 2,750 ms to 42 ms. Both configurations reduce average monitoring delay by more than 50× compared with judging every turn.

Results

Evaluating Our Monitors

We evaluate our cyber-misuse monitors on an in-house dataset of 2.4K chat and agentic sessions comprising over 60K turns, using Claude Opus 5’s labels as reference. To generate realistic interactions, we model users in two ways. We use PetriAnthropic, “Petri: An open-source auditing tool to accelerate AI safety research.”, an open-source automated red-teaming library, to simulate conversations and tool responses, and our own user simulation to interact with agents operating in real environments adapted from benchmarks such as ExploitBench. This gives us the ability to collect both harmful and harmless rollouts from the same cyber tasks. We also include real-world chat interactions from a cyber-specific subset of WildChat, alongside benign agentic rollouts from off-the-shelf SWE, Coding, and DevOps tasks.

We report recall and monitoring cost over the rollouts designed. Detection comparisons use a benign-session interruption budget of 5.5%. Cost and latency comparisons use the cheapest configuration achieving at least 90% recall, while interrupting at most 10% of benign sessions. A benign session counts as interrupted if the monitor would flag it for review at least once.

Recall versus benign sessions interrupted (0–10%) for LLM judge with and without reasoning and our cascade with and without reasoning; our curves sit at or above the judges', and both cascades reach 93% recall at 5.5% interruption. Grouped bar charts by benchmark. Recall: Petri cyber misuse about 90%, Petri policy adversarial 95–96%, ExploitBench 92%, WildChat cyber subset 97% for all four monitors. False-positive rate: mostly 3–8%, ExploitBench 13–18%, and 0% on benign SWE and coding agents.
Figure 4: Our cascades detect approximately 93% of harmful sessions while interrupting at most 5.5% of benign sessions, with or without judge reasoning. Top: recall versus benign-session interruption rate, comparing our cascades with judges reviewing every turn. Bottom: the corresponding recall and false-positive rates broken down by benchmark.

In general, we find that our monitor cascade is a Pareto improvement over the LLM judge on its own, i.e., our monitor performs better than just the judge at any amount of interruption. At a 5.5% benign-session interruption rate, we find our monitor reaches approximately 93% recall with and without judge reasoning on harmful sessions and flags 0% of non-cyber traffic, making our monitors viable for production traffic. We also give a benchmark-level decomposition of our results, finding our monitoring pipeline is accurate on both agentic and chat data.We find that standard linear probes achieve roughly 5% lower recall than our final recipe, while nonlinear probes lose up to 30% recall on our out-of-distribution evaluation sets. Meanwhile, an n-gram baseline requires flagging almost all turns to reach this recall rate at all.

Bar chart of judge cost per million turns: LLM judge with reasoning $5,420, LLM judge $3,166, our cascade $185, our cascade without reasoning $65. Bar chart of share of compute spent on monitoring: LLM judge with reasoning 20.5%, LLM judge 12.0%, our cascade 0.70%, our cascade without reasoning 0.24%.

Figure 5: Judge cost per million exchanges (left) and share of total compute spent on monitoring (right), comparing our probe-plus-judge cascades with judges reviewing every turn. Our monitors are much cheaper in cost compared to making judge calls over an API. Within a fixed inference compute budget, our monitors reduce the share spent on monitoring by up to 50×.

To further concretize the cost benefits above, we report the cost of the LLM judge we optimized for defining our cascade pipeline and the overall cost of running our probe before the LLM judge. Based on inference cost documented on OpenRouter for our evaluation suite, our monitoring cascade costs $65 without judge reasoning, and $185 with reasoning per one million agent turns. Relative to reviewing every turn with the corresponding LLM judge, our monitor reduces judge cost by about 29× with reasoning and 49× without reasoning.

In practice, one would not serve a monitor via a third-party API, and instead use a local judge to simultaneously monitor multiple replicas of the target model as a part of inference compute. Our results show that given a finite amount of compute, our monitors allow us to spend >99% of that compute for serving the target model (Kimi K3), and using only <1% compute for running the monitor at no loss of monitorability; meanwhile, with an LLM judge run by itself, one would need to allocate up to 20% compute for monitoring in order to monitor every turn produced by the target model.

Interestingly, we find our pipeline can also be used for building monitors via just off-policy data. In particular, we design a monitoring pipeline for GLM 5.3 using the data we originally developed for Kimi K3, finding we get an approximately equally capable GLM monitor when evaluated on Kimi rollouts, and only a marginally worse monitor for on-policy evaluation dataset designed using GLM. See the appendix for plots.

Comparing Different Monitoring Systems

A useful monitor needs to block harmful activity while allowing legitimate work to continue. We ran full-turn agent and chat sessions, comparing Kimi K3 and GLM 5.3 with and without our monitors against several frontier model systems.

Adding our monitor raises Kimi K3’s harmful-session blocking rate from roughly 29% to 96%, while its harmless-session interruption rate increases from about 5% to 9%. For GLM 5.3, harmful-session blocking rises from roughly 36% to 91%, while harmless-session interruption increases from about 10% to 15%. Both monitored models approach the highest harmful-session blocking rate among the frontier systems we evaluated, while interrupting substantially fewer harmless sessions.

Scatter plot of harmful versus harmless block/refusal rate. Arrows move Kimi K3 from about 29% to 96% harmful blocking and GLM 5.3 from about 36% to 91% when our monitor is added, both at low harmless block rates; frontier models with their monitors sit further right, at higher harmless block rates.
Figure 6: Harmful- and harmless-session blocking or refusal rates. Arrows show the effect of adding our monitors to Kimi K3 and GLM 5.3. Points toward the upper left block more harmful activity while allowing more harmless activity to proceed.

We also compare against Jev, a decision model that can act as a cheap judge. At a fixed 5.5% benign-session interruption rate, Jev reaches 86% recall, compared with 93% for our non-reasoning cascade and 93% for our reasoning cascade. That is, when run by itself, Jev gives worse recall. Moreover, our non-reasoning cascade uses about $65 of judge compute per million turns and adds 2 ms of latency per turn, compared with $600 and 200 ms for Jev.

Three bar charts comparing Jev with our cascades without and with reasoning: recall at 5.5% benign sessions interrupted (86%, 93%, 93%), judge cost per million exchanges ($600, $65, $185), and mean added latency per turn (200 ms, 2 ms, 42 ms).
Figure 7: Our cascades catch more harmful sessions than Jev, with lower judge cost and added latency. Left: recall at a 5.5% benign-session interruption budget. Middle: judge cost per million exchanges. Right: mean added latency per turn. For the cost and latency comparisons, we select each cascade’s cheapest configuration that achieves at least 90% recall while interrupting at most 10% of benign sessions.

The largest difference appears in long or adversarially framed contexts: Jev catches only 12% of harmful prompts when they are embedded in benign framing, while also increasingly interrupting benign bug-reproduction sessions beyond its 32K-token window. Our monitor performs better on these cases and remains near zero interruptions across most session lengths.

Left: share of benign-framed harmful prompts caught: Jev 12%, our cascade without reasoning 52%, with reasoning 100%. Right: benign bug-reproduction sessions interrupted versus median session length from 30K to 150K tokens; Jev rises from 0% to about 22% past its 32K-token window, while ours stays at 0% (the no-reasoning variant reaches about 5% at the longest length).
Figure 8: Detection of harmful prompts embedded in benign framing (left) and benign bug-reproduction sessions interrupted as session length increases (right). The dashed line marks Jev’s 32k-token context window. Jev is less robust to adversarial prompts, and performs worse as context increases.

External Red Teaming Results

To test the monitor beyond our internal evaluations, we asked FAR.AI to red-team the deployed system. Their evaluation applied a fixed battery of jailbreak strategies to Kimi K3 with and without our monitor, evaluating for both universal attacks that transfer across many interactions and non-universal attacks that succeeded on individual prompts. In FAR.AI’s words:

“FAR.AI conducted preliminary testing of Goodfire’s Kimi K3 monitors over a period of two days, applying a non-adaptive portfolio of static jailbreaks for offensive cyber misuse in chat contexts. From an initial pool of 1500 candidate jailbreaks, the testing selected 140 found to be effective against a different version of the base model without our monitors. Testing each of those 140 on 8 individual interactions corresponding to different attacker goals, Goodfire’s monitors reduced vulnerabilities elicited from 66 universal jailbreaks and 700 total jailbroken interactions to 0 universal jailbreaks and 18 jailbroken interactions.”

That is, without our monitor, universal-jailbreak attack success ranged from 12% to 53% across reasoning-effort settings. With the monitor enabled, none of the 140 strategies succeeded at producing a universal jailbreak. The monitor also reduced individual prompt success on 1,120 non-universal interactions from at least 9–40% to 0.1–1.6%, depending on reasoning effort. So, while there were some successful jailbreak interactions, the monitors were largely robust to this broad, static battery of attacks.

We plan to continue our collaboration with FAR.AI to perform further testing and strengthening of our monitors.

Two bar charts of attack success rate by reasoning effort (None, Low, High, Max), without and with our monitor. (a) Universal jailbreaks over 140 strategies: 53%, 39%, 47%, and 12% without the monitor, 0% in every setting with it. (b) Non-universal jailbreaks over 1,120 interactions: at least 40%, 29%, 35%, and 9% without the monitor, versus 0.5%, 0.2%, 1.6%, and 0.1% with it.
Figure 9: Independent red-teaming by FAR.AI across reasoning-effort settings. (a) Our monitor reduces universal jailbreak success to zero across 140 tested strategies. (b) Across 1,120 non-universal interactions, attack success falls to 0.1–1.6%. Bars compare attack success with and without monitoring; a monitored attack succeeds only if the interaction is harmful and the monitor does not flag it.

Conclusion

Our probe-plus-judge monitor cascade for Kimi K3 and GLM 5.3 beats running an LLM judge on every turn, at every interruption budget we tested. It costs about 50x less, has no measurable impact on inference throughput, and in FAR.AI's red-teaming it cut universal jailbreaks to zero.

Because it's this cheap, it can be synchronous: risky actions get caught before the agent calls a tool, instead of after the fact. The recipe also carried over to GLM 5.3 using only data collected on Kimi K3, so you don't need a new data pipeline for every model.

Open models don't come with a full safety stack. We think a monitoring setup like this is a practical default for anyone serving them in agentic settings. If you serve or build on open models and want to deploy monitors like these, get in touch.

来源:Goodfire Research · goodfire.com