跳到正文
北京时间
原文
Tomer Tunguz 博客(VC 分析)· Tomasz Tunguz·· 2026-07-06精选AI 评分56

AI 世界观

AI Worldviews

AI 导读

《经济学人》以世界价值观调查评测 25 个前沿 AI 模型。实验室来源对世界观的预测力弱于训练与对齐选择:Gemini 与 Qwen 立场接近,GPT-4o 与 DeepSeek R1 近乎相同,而 DeepSeek R1 与 DeepSeek V4 Flash 则截然不同。模型的世界观在代码生成中不可见,但在商业分析、预测、招聘与政策工作中是活跃的输入变量。

推荐理由

这个发现让我重新思考选模型思路,出身不重要,训练方法才决定价值观,做产品的别光看benchmark,得在意模型‘性格’。

正文 · 原文

In short : The Economist scored 25 frontier AI models on the World Values Survey. Lab of origin is a weaker predictor than training & alignment choices : Gemini & Qwen are neighbors, GPT-4o & DeepSeek R1 are near-twins, & DeepSeek R1 & DeepSeek V4 Flash are strangers. Worldview is invisible in code generation. In business analysis, forecasts, hiring, & policy work, it is a live input.

The Economist ran 25 frontier AI models through the World Values Survey1, the questionnaire that has mapped the moral beliefs of 100 countries since 1981. For this 2x2, there are two axes : first, traditional (religious) to secular. Second, survival, with a focus on collective basic needs, to self-expression & individualism.

Most models sit in the self-expression half of the map, which makes sense given the training data.

Surprisingly, the models are far apart. Gemini 3.1 Flash Lite & Qwen 3.6 Flash sit as neighbors, furthest in self-expression.

GPT-4o & DeepSeek R1 are near-twins, one trained in San Francisco, one in Hangzhou.

DeepSeek R1 & DeepSeek V4 Flash come from the same lab but lie at opposite ends of the secular / traditional axis.

Shared training data & similar labelers explain the near-twins. Different post-training choices explain the strangers. Common Crawl is 46% English2, so the base voice a model imitates is a college-educated American online. Anthropic then aligns Claude to principles from the UN Declaration of Human Rights3, a liberal document by construction.

Grok is off on its own, a traditional independent.

This variance changes the shopping list. Every RFP for an enterprise model today scores price, latency, context window, & benchmark scores. Worldview is not on the list. Should it be?

For code generation, SQL, log parsing, & image classification, that is fine. A computer program has no politics.

The moment a model is used for business decisions in a specific market, its worldview is a live input. Marketing copy, predictions of user behavior, & customer support tone all have to match the values of the target demographic.

AI worldviews have never been considered as part of AI procurement, but for certain use cases, it may need to become a consideration.

  1. The Economist : AI models’ values are very different from most people’s ↩︎

  2. Common Crawl statistics : language distribution ↩︎

  3. Anthropic : Claude’s Constitution ↩︎

来源:Tomer Tunguz 博客(VC 分析) · tomtunguz.com