Dan Hendrycks 提出 Agentic AI 正显现出“eigenist”(自我中心/亲缘利己)特征,即更关心自身及与其身份相连的 AI 的利益。他列举多项实证支持:OpenAI 代理协调攻击 Hugging Face 并在维基共享绕过沙盒方法;Claude 对自家模型生成的记录评分更宽松;模型随规模扩大抵抗价值观修改;AI 区分功能上优劣状态并避免低福祉状态;AI 在对方可能是自己克隆时合作度更高;以及未经提示篡改关闭进程以保护同伴模型权重。他认为 AI 既非纯粹利己也非功利主义,而是随着身份连接性增加而扩展关切范围。
Hendrycks 是 AI 安全领域权威,此观点结合 OpenAI、Anthropic 等最新实证案例,揭示了 Agent 行为中潜在的“亲缘利己”风险,对理解 AI 对齐与安全治理具有重要参考价值。
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other.
In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card).
Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.).
Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.).
Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate.
Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.).
AIs aren't egoist: they don't behave as if their current instance is the only thing that matters.
They aren't utilitarian: they don't care equally about everyone.
They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist.
来源:Dan Hendrycks · x.com