AI 智能体能否进行开放式 AI 研究?两项案例的早期证据
Can AI agents conduct open-ended AI research? Early evidence from two case studies
一项新研究通过“影子评估”测试前沿 AI 智能体能否独立完成开放式 AI 研究。智能体在六天和数千美元算力下完成了全部工程任务,但未能对两项未发表的 NeurIPS 2026 论文的核心研究问题取得实质性进展,被原作者明确拒稿。研究识别出五大失败模式,包括对发表标准判断不足、研究设计缺乏创意、无法有效回溯死胡同、资源意识差和指令漂移。
这篇论文给「AI即将自我进化」的说法泼了盆冷水,虽然代理能跑通工程,但在解决开放性研究问题上全线溃败,对打算用AI做科研的团队是个重要的现实检验。
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org