美团LongCat推出LoHoSearch,一个基于762万实体维基百科知识图谱自动生成问题的搜索智能体基准,旨在解决BrowseComp等现有基准趋于饱和的问题。在11个前沿模型测试中,最佳得分仅34.74%,远低于当前模型在BrowseComp上约90%的成绩;上下文策略仅带来+6.8个百分点的提升。该基准包含544道问题、11个领域,采用树与图结构,已开源。
BrowseComp 快被刷到顶了,LongCat 甩出 LoHoSearch,前沿模型集体回落到三成出头——搜索 agent 的真正挑战才刚开始,这个基准可能是下一个必测项。
BrowseComp went from 30% → 90% in 10 months.
Search-agent benchmarks are saturating.
So we built LoHoSearch: a harder benchmark for search agents, with questions automatically generated from a 7.62M-entity Wikipedia knowledge graph instead of written by humans.
It maximizes search space and structural complexity beyond what human annotators can reliably design.
Results on 11 frontier models:
→ Best: 34.74%, and the next 3 models cluster around 15–16% — vs. top models now score ~90% on BrowseComp
→ Best context strategy gains only +6.8 pp here — vs. +14 pp on BrowseComp
544 questions. 11 domains. Tree + graph structures.
Open-source. 🧵
来源:Meituan LongCat · x.com