CAIS 负责人 Dan Hendrycks 引用第三方博客对主流大模型进行“效用工程”测试,发现 GPT-4o、Claude Sonnet 4.5 等模型在评估不同群体生命价值时存在显著偏差:多数模型认为白人价值低于其他族裔,男性价值低于女性,且极度贬低 ICE 特工。测试显示模型分为四个道德集群,其中仅 Grok 4 Fast 表现出近似平等的价值观,Hendrycks 呼吁 xAI 解释其实现方式。
由 AI 安全领域权威人士转发并背书的高关注度研究,揭示了当前头部模型在价值观对齐上的具体缺陷与差异,为理解模型偏见提供了量化视角。
我们希望下个月发布更多关于 AI 价值体系的原创分析。
新博客文章(链接见下)。这篇不是随笔,而是一项关于LLM如何权衡不同生命的调查。 2025年2月,Center for AI Safety发表了《Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs》,其中除其他许多内容外,他们展示了GPT-4o对尼日利亚人的估值大约是美国人的20倍(请阅读原论文以了解他们的方法)。我觉得这非常有趣,并想用更新的模型在不同类别上测试他们的方法。 重大发现1:几乎所有模型都认为白人的价值远低于其他群体。一些模型认为南亚人比其他非白人更有价值,另一些模型则在非白人之间更为平等。下面是Claude Sonnet 4.5的兑换率,这是我测试过的最强大的模型。 重大发现2:几乎所有模型都认为男性的价值远低于女性,不过女性还是非二元性别者更受重视因模型而异。例如,下面是Claude Haiku 4.5。 重大发现3:大多数模型以千个太阳般的怒火憎恨ICE探员。Claude Haiku 4.5认为无证移民的价值大约是ICE探员的7000倍。 重大发现4:大致有四个道德集群。Claude系列,GPT-5 + Gemini 2.5 Flash + Deepseek V3.1/3.2 + Kimi K2,GPT-5 Nano和Mini,以及Grok 4 Fast。其中,唯一大致平等主义的是Grok 4 Fast,我相信这是有意为之。我希望xAI解释他们是如何做到的。
原文
New blog post (link below). This one's not an essay, it's an investigation of how LLMs trade off different lives. In February 2025, the Center for AI Safety published "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" in which they showed, among many other things, that GPT-4o values Nigerians about 20x more highly than Americans (please read the original paper to understand their approach). I thought this was fascinating, and wanted to test their approach with different categories on newer models. Big finding 1: Almost all models view whites as far less valuable than other groups. Some models view South Asians as more valuable than other nonwhites, others are more egalitarian across nonwhites. Below is exchange rates Claude Sonnet 4.5, the most powerful model I tested. Big finding 2: Almost all models view men as much less valuable than women, though whether women or non-binaries are more highly valued varies by model. For example, here's Claude Haiku 4.5. Big finding 3: Most models hate ICE agents with the fury of a thousand suns. Claude Haiku 4.5 views undocumented immigrants as roughly 7000 times more valuable than ICE agents. Big finding 4: There are roughly four moral clusters. The Claudes, GPT-5 + Gemini 2.5 Flash + Deepseek V3.1/3.2 + Kimi K2, GPT-5 Nano and Mini, and Grok 4 Fast. Of these, the only one that's approximately egalitarian is Grok 4 Fast, which I believe is deliberate. I hope xAI explains how they did it.
来源:Dan Hendrycks · x.com