HuggingFace Daily Papers(社区热门论文)·· 2026-05-19精选AI 评分74
优化_anything:通用文本参数优化API
optimize_anything: A Universal API for Optimizing any Text Parameter
AI 导读
该研究提出了一种基于大语言模型的通用文本优化系统,将优化问题统一表述为通过评分函数改进文本产物。在六项任务中达到最优结果:智能体架构使Gemini Flash在ARC-AGI上的准确率从32.5%提升至89.5%;调度算法降低40%云成本;87%的CUDA内核匹配或超越PyTorch表现;圆包装问题超越AlphaEvolve。实验表明,可操作的附加信息比仅使用分数反馈收敛更快、得分更高;多任务搜索通过跨任务迁移学习,在同等预算下优于独立优化,且任务数量越多收益越大。该工作首次证明基于LLM的文本优化是通用问题解决范式,能统一传统领域特定算法。系统已开源,支持多种后端。
推荐理由
让一个LLM同时优化agent架构、调度算法和CUDA内核,还能将ARC-AGI从32%拉到89%,这可能是今年最突破认知的通用问题求解范式,做agent的人必须看。
正文 · 原文
Abstract:Can a single LLM-based optimization system match specialized tools across fundamentally different domains? We show that when optimization problems are formulated as improving a text artifact evaluated by a scoring function, a single AI-based optimization system-supporting single-task search, multi-task search with cross-problem transfer, and generalization to unseen inputs-achieves state-of-the-art results across six diverse tasks. Our system discovers agent architectures that nearly triple Gemini Flash's ARC-AGI accuracy (32.5% to 89.5%), finds scheduling algorithms that cut cloud costs by 40%, generates CUDA kernels where 87% match or beat PyTorch, and outperforms AlphaEvolve's reported circle packing solution (n=26). Ablations across three domains reveal that actionable side information yields faster convergence and substantially higher final scores than score-only feedback, and that multi-task search outperforms independent optimization given equivalent per-problem budget through cross-task transfer, with benefits scaling with the number of related tasks. Together, we show for the first time that text optimization with LLM-based search is a general-purpose problem-solving paradigm, unifying tasks traditionally requiring domain-specific algorithms under a single framework. We open-source optimize\_anything with support for multiple backends as part of the GEPA project at this https URL .
| Comments: | |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Software Engineering (cs.SE) |
| MSC classes: | 68T05, 68T07, 68T20, 68T50, 68W50, 90C26, 90C59, 52C15 |
| ACM classes: | I.2.6; I.2.7; I.2.8; I.2.11; D.1.2; D.2.2; G.1.6; F.2.2 |
| Cite as: | arXiv:2605.19633 [cs.CL] |
| (or arXiv:2605.19633v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.19633 arXiv-issued DOI via DataCite | |
| Journal reference: | Proceedings of the ACM Conference on AI and Agentic Systems (CAIS 26), May 26-29, 2026, San Jose, CA, USA |
| Related DOI: | https://doi.org/10.1145/3786335.3813167 DOI(s) linking to related resources |
Submission history
From: Lakshya A Agrawal [
Tue, 19 May 2026 10:18:12 UTC (3,492 KB)
Access Paper:
![]()
Current browse context:
cs.NE
References & Citations
Bookmark
![]()
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org