跳到正文
北京时间
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-05-13精选AI 评分72

AgentLens:揭示软件工程智能体评估中的“幸运通过”问题

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

AI 导读

当前软件工程智能体评估仅依赖最终补丁是否通过测试的二元信号,掩盖了解决方案质量的差异。研究分析了2,614条轨迹,发现在可评估的1,815条通过轨迹中,10.7%属于“幸运通过”,表现为回归循环、盲目重试等问题。为此,研究团队提出了用于过程级评估的AgentLens框架,并发布了标注质量分数、冗余信号等信息的AgentLens-Bench数据集。基于质量分数,通过轨迹被划分为幸运、扎实和理想三个等级,不同模型的幸运通过率介于0.5%至23.2%之间。若按质量分数而非通过率排名,部分模型的排名变化显著。相关资源已开源。

推荐理由

SWE-agent评估只看通过率太粗暴了,这篇论文把乱试的“幸运通过”和真方案拆开看,10%的通过其实是蒙的,做agent评估的必读。

正文 · 原文

Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on 60 SWE-bench Verified tasks. Of these, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and define AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens builds PTA references by merging multiple passing solutions for the same task, and uses a context-sensitive intent labeler to assign actions to Exploration, Implementation, Verification, or Orchestration based on trajectory history rather than tool identity alone. On AgentLens-Bench, the quality score separates passing trajectories into Lucky, Solid, and Ideal tiers and further decomposes Lucky Passes into five recurring mechanisms. Across the eight model backends, Lucky rates range from 0.5% to 23.2%, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We plan to release the project repository soon, including AgentLens-Bench artifacts, the AgentLens SDK, and the analysis tooling.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org