跳到正文
北京时间
原文
Meituan LongCat· @Meituan_LongCat · X·· 2026-07-17精选AI 评分75
AI 导读

美团LongCat推出LoHoSearch,一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准,旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中,最佳得分仅34.74%,远低于当前模型在BrowseComp上约90%的成绩;上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域,采用树与图结构,已开源。

推荐理由

BrowseComp 快被刷到顶了,LongCat 甩出 LoHoSearch,前沿模型集体回落到三成出头——搜索 agent 的真正挑战才刚开始,这个基准可能是下一个必测项。

正文 · 原文

BrowseComp went from 30% → 90% in 10 months.
Search-agent benchmarks are saturating.

So we built LoHoSearch: a harder benchmark for search agents, with questions automatically generated from a 7.62M-entity Wikipedia knowledge graph instead of written by humans.
It maximizes search space and structural complexity beyond what human annotators can reliably design.

Results on 11 frontier models:
→ Best: 34.74%, and the next 3 models cluster around 15–16% — vs. top models now score ~90% on BrowseComp
→ Best context strategy gains only +6.8 pp here — vs. +14 pp on BrowseComp

544 questions. 11 domains. Tree + graph structures.
Open-source. 🧵

来源:Meituan LongCat · x.com