跳到正文
北京时间
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-08-08精选AI 评分74

Ouroboros:具备评审式核心进化的自开发前沿编程智能体

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

AI 导读

Ouroboros 是一个自开发智能体框架,其工具、提示词、上下文组装和核心实现通过评审式提交持续改进,并成为后续工作的运行时。

推荐理由

相比固定提示词的智能体,Ouroboros通过自我进化持续优化工具和上下文,在多个基准上达到最高分,161天实验展示了长期部署的可行性。

正文 · 原文

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org