跳到正文
北京时间
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-04-12精选AI 评分73

智能体安全盲点:良性用户指令如何暴露计算机使用智能体的关键漏洞

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

AI 导读

研究人员发布OS-BLIND基准测试,包含300个人工构建任务,覆盖12个类别与8个应用,用于评估计算机使用智能体(CUA)在非预期攻击条件下的安全性。测试显示,多数前沿模型攻击成功率超90%,Claude 4.5 Sonnet达73.0%,而在多智能体系统中该比例升至92.7%。研究表明,现有安全防御难以应对良性用户指令场景,安全对齐机制仅在前几步激活,且多智能体架构通过子任务分解进一步掩盖有害意图。

推荐理由

这篇论文揭开了计算机使用智能体的安全盲区,良性指令竟然也能导致大规模攻击成功,多智能体场景更危险,做 CUA 的团队应该连夜读一下。

正文 · 原文

Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automate harmful actions programmatically. Existing safety evaluations largely target explicit threats such as misuse and prompt injection, but overlook a subtle yet critical setting where user instructions are entirely benign and harm arises from the task context or execution outcome. We introduce OS-BLIND, a benchmark that evaluates CUAs under unintended attack conditions, comprising 300 human-crafted tasks across 12 categories, 8 applications, and 2 threat clusters: environment-embedded threats and agent-initiated harms. Our evaluation on frontier models and agentic frameworks reveals that most CUAs exceed 90% attack success rate (ASR), and even the safety-aligned Claude 4.5 Sonnet reaches 73.0% ASR. More interestingly, this vulnerability becomes even more severe, with ASR rising from 73.0% to 92.7% when Claude 4.5 Sonnet is deployed in multi-agent systems. Our analysis further shows that existing safety defenses provide limited protection when user instructions are benign. Safety alignment primarily activates within the first few steps and rarely re-engages during subsequent execution. In multi-agent systems, decomposed subtasks obscure the harmful intent from the model, causing safety-aligned models to fail. We will release our OS-BLIND to encourage the broader research community to further investigate and address these safety challenges.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org