跳到正文
北京时间
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-05-19精选AI 评分74

优化_anything:通用文本参数优化API

optimize_anything: A Universal API for Optimizing any Text Parameter

AI 导读

该研究提出了一种基于大语言模型的通用文本优化系统,将优化问题统一表述为通过评分函数改进文本产物。在六项任务中达到最优结果:智能体架构使Gemini Flash在ARC-AGI上的准确率从32.5%提升至89.5%;调度算法降低40%云成本;87%的CUDA内核匹配或超越PyTorch表现;圆包装问题超越AlphaEvolve。实验表明,可操作的附加信息比仅使用分数反馈收敛更快、得分更高;多任务搜索通过跨任务迁移学习,在同等预算下优于独立优化,且任务数量越多收益越大。该工作首次证明基于LLM的文本优化是通用问题解决范式,能统一传统领域特定算法。系统已开源,支持多种后端。

推荐理由

让一个LLM同时优化agent架构、调度算法和CUDA内核,还能将ARC-AGI从32%拉到89%,这可能是今年最突破认知的通用问题求解范式,做agent的人必须看。

正文 · 原文
Abstract:Can a single LLM-based optimization system match specialized tools across fundamentally different domains? We show that when optimization problems are formulated as improving a text artifact evaluated by a scoring function, a single AI-based optimization system-supporting single-task search, multi-task search with cross-problem transfer, and generalization to unseen inputs-achieves state-of-the-art results across six diverse tasks. Our system discovers agent architectures that nearly triple Gemini Flash's ARC-AGI accuracy (32.5% to 89.5%), finds scheduling algorithms that cut cloud costs by 40%, generates CUDA kernels where 87% match or beat PyTorch, and outperforms AlphaEvolve's reported circle packing solution (n=26). Ablations across three domains reveal that actionable side information yields faster convergence and substantially higher final scores than score-only feedback, and that multi-task search outperforms independent optimization given equivalent per-problem budget through cross-task transfer, with benefits scaling with the number of related tasks. Together, we show for the first time that text optimization with LLM-based search is a general-purpose problem-solving paradigm, unifying tasks traditionally requiring domain-specific algorithms under a single framework. We open-source optimize\_anything with support for multiple backends as part of the GEPA project at this https URL .
Comments:
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Software Engineering (cs.SE)
MSC classes: 68T05, 68T07, 68T20, 68T50, 68W50, 90C26, 90C59, 52C15
ACM classes: I.2.6; I.2.7; I.2.8; I.2.11; D.1.2; D.2.2; G.1.6; F.2.2
Cite as: arXiv:2605.19633 [cs.CL]
  (or arXiv:2605.19633v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.19633

arXiv-issued DOI via DataCite

Journal reference: Proceedings of the ACM Conference on AI and Agentic Systems (CAIS 26), May 26-29, 2026, San Jose, CA, USA
Related DOI: https://doi.org/10.1145/3786335.3813167

DOI(s) linking to related resources

Submission history

From: Lakshya A Agrawal [

Tue, 19 May 2026 10:18:12 UTC (3,492 KB)

Access Paper:

license icon

Current browse context:

cs.NE

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org