跳到正文
北京时间
原文
OpenBMB· @OpenBMB · X·· 2026-07-31精选AI 评分75
AI 导读

面壁智能与清华NLP团队提出ALIGN,自动生成对齐接口解决智能体与环境间的失配问题。仅改写反馈措辞即可将Qwen2.5-7B智能体在ALFWorld上的成功率从13.4%提升至31.3%。该方法在四个基准上最高提升45.67%成功率,并减少65%连续无效动作,且接口可跨智能体架构和LLM骨干迁移。

推荐理由

做 agent 的人都知道环境对齐是个隐形的坑,这篇论文不仅把问题讲透了,还给了开箱即用的 wrapper,ALFWorld 上成功率从 13% 拉到 31%,原因只是改写了反馈措辞,我觉得这是今年 agent 工程最被低估的发现。

正文 · 原文

Everyone building LLM agents pours effort into the agent's strategy or into harder environments. The interface between them gets almost no attention, and it is often where agents actually break.

Case in point in ALFWorld: an agent tries examine shelf 1, but the env requires go to first, so it just returns "Nothing happens", and the agent concludes the shelf is empty. Simply rewording that feedback lifts a Qwen2.5-7B agent from 13.4% to 31.3%.

ALIGN from @TsinghuaNLP (OpenBMB member) automatically generates aligned interfaces to fix this misalignment.
Paper: https://arxiv.org/abs/2505.21055
Code: https://github.com/THUNLP-MT/ALIGN

1⃣️ The problem is agent-environment misalignment: the agent's expectation of what an action does diverges from the environment's real transitions, because implicit rules and under-specified observations are never surfaced. The paper shows this is a pervasive bottleneck, not an agent reasoning failure.

2⃣️ ALIGN wraps the environment with two modules. INFERRULES surfaces static rules and constraints (preconditions, action ordering) before the task; WRAPSTEP intercepts each action and enriches the raw observation with success/failure conditions. It is a lightweight Python wrapper, no changes to agent logic or environment code.

3⃣️ Interfaces are generated iteratively by an Analyzer (diagnoses misalignments from failed trajectories) and an Optimizer (synthesizes and refines the interface as Python functions). Both run experimental verification against the live environment to fight hallucination; ablating it collapses accuracy.

4⃣️ Across four benchmarks in embodied, web, and tool-use, gains reach +45.67% success on ALFWorld and cut consecutive invalid actions by 65%. Interfaces transfer plug-and-play across agent architectures (ReAct, Self-Consistency, Planning...) and across LLM backbones (Qwen, Llama) with no regeneration.

来源:OpenBMB · x.com