跳到正文
北京时间
原文
Anthropic· @AnthropicAI · X·· 2025-10-07精选
AI 导读

Anthropic 上周发布 Claude Sonnet 4.5,期间使用新工具对模型进行自动化对齐审计以检测谄媚与欺骗行为。该工具现已开源。

推荐理由

Anthropic 开源对齐测试工具,可审计模型谄媚与欺骗行为

正文 · 原文

Last week we released Claude Sonnet 4.5. As part of our alignment testing, we used a new tool to run automated audits for behaviors like sycophancy and deception.

Now we’re open-sourcing the tool to run those audits. https://t.co/cCJGNaVFrl

来源:Anthropic · x.com