Weblica:面向视觉网页智能体的可扩展可复现训练环境
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
苹果研究团队提出Weblica框架,通过HTTP级缓存保存网页稳定视觉状态并保留交互行为,结合大语言模型基于真实网站与核心导航技能合成环境,构建可复现、可扩展的训练环境。该框架将强化学习训练扩展到数千个多样化的环境和任务。最佳模型Weblica-8B在多个网页导航基准上超越同等规模的开源模型,推理步骤更少,测试时计算扩展性良好,性能与API模型相当。
Apple 的研究把 web agent 训练从零散数据中解放出来,可复现环境是规模化 RL 的关键一步,做 web 自动化的值得关注。
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stable visual states while preserving interactive behavior and 2) LLM-based environment synthesis grounded in real-world websites and core web navigation skills. Using this framework, we scale RL training to thousands of diverse environments and tasks. Our best model, Weblica-8B, outperforms open-weight baselines of similar size across multiple web navigation benchmarks while using fewer inference steps, scales favorably with additional test-time compute, and is competitive with API models.
来源:Apple Machine Learning Research(RSS) · machinelearning.apple.com