跳到正文
北京时间
原文
Kimi.ai· @Kimi_Moonshot · X·· 2026-07-28精选AI 评分73
AI 导读

Kimi.ai 发布 PerceptionBench,一个从当前前沿模型在 42 个基准上的失败模式中归纳出的视觉感知基准。该基准将视觉感知拆解为 10 种原子能力,并构建了 3000 道验证题,每道题只考察单一感知能力,无需推理或外部知识。

推荐理由

从模型失败中反向定义原子感知,把视觉感知从推理里拆出来很聪明,开源数据和方法对做视觉评测的产品人很有参考价值。

正文 · 原文

We are releasing PerceptionBench, a benchmark that isolates visual perception and evaluates it as a set of atomic capabilities - discovered from how today's models fail, rather than defined in advance.

From frontier-model failures across 42 benchmarks, we derive 10 atomic perceptual capabilities and construct 3,000 verified questions, each isolating a single capability and answerable by looking, with no reasoning or external knowledge required.

Blog: http://kimi.com/blog/perception-bench
GitHub: http://github.com/MoonshotAI/PerceptionBench
Hugingface: http://huggingface.co/datasets/moonshotai/PerceptionBench

来源:Kimi.ai · x.com