跳到正文
北京时间
原文
Dongxi 东锡 NLP· @dongxi_nlp · X·· 2 小时前AI 评分46
AI 导读

You Only Edit Once: 通过局部示例精修激发 LLM 的上下文能力 挑选最佳少样本示例是一种缓慢的 System-2 搜索:组合爆炸,且往往需要反复调用 LLM。 我们把它变成了 System 1。⚡ Jev-LDE,一个 1.7B 的编辑器,扫一眼检索到的示例,只做一次编辑。LLM 只回答一次。 平均 1-shot 准确率 81.2 → 88.1 You Only Edit Once 🧵 动画演示(示意示例)。查询:“How far is it from Denver to Aspen?” 语义 TopK 检索出三个相似示例:“Where is Aspen, Colorado?”(Location)、“What state is Denver in?”(Location)、“Who founded Denver?”(Person)。Jev-LDE,一个 1.7B 的 System-1 编辑器,标记出第一个虽然主题相同但答案类型错误,并输出一个动作:将 S1 替换为候选 C1,“How far is Boston from NYC?”(Number)。冻结的目标 LLM 随后回答“Number”,这是正确的;若不编辑,它会回答“Location”。结尾卡片:在 3 个基准和 4 个目标 LLM 上,平均 1-shot 准确率从 81.2 提升至 88.1,在 48 个设置中有 44 个达到最佳或并列最佳,墙钟时间增加 11%。

正文

You Only Edit Once:

Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

引用Dr. Cheems Wang 🏡@AlbertW24045555
You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement Picking the best few-shot demos is a slow System-2 search: combinatorial, often with repeated LLM calls. We made it System 1. ⚡ Jev-LDE, a 1.7B editor, glances at the retrieved demos and makes ONE edit. The LLM answers once. Avg 1-shot acc 81.2 → 88.1 You Only Edit Once 🧵 Animated walkthrough (illustrative example). Query: "How far is it from Denver to Aspen?" Semantic TopK retrieves three look-alike demos: "Where is Aspen, Colorado?" (Location), "What state is Denver in?" (Location), "Who founded Denver?" (Person). Jev-LDE, a 1.7B System-1 editor, flags the first as same topic but wrong answer type and outputs one action: Replace S1 with candidate C1, "How far is Boston from NYC?" (Number). The frozen target LLM then answers "Number", which is correct; without the edit it answers "Location". End card: average 1-shot accuracy 81.2 to 88.1 across 3 benchmarks and 4 target LLMs, best or tied-best in 44 of 48 settings, +11% wall time.
在 X 查看被引用的帖子

来源:Dongxi 东锡 NLP · x.com