Berkeley RDI 发布开源平台 CUA-Lite,面向计算机使用智能体
CUA-Lite: An Open Platform for Computer-Use Agents
Berkeley RDI 推出开源平台 CUA-Lite,为计算机使用智能体提供三个标准化抽象。其环境接口下运行 15+ 个覆盖桌面、浏览器和移动端的基准,并提供含 30k+ 可验证任务的免 VM 桌面沙箱;统一监督数据格式已转换 10+ 个公开 CUA 数据集;每模型一个 harness 在评测、SFT 和 RL 间共享,支持 14 个模型族。
原文说明了平台的三项标准化抽象及覆盖规模,读者可以据此评估它在整合分散的计算机使用智能体资源上的可用性。
Introducing CUA-Lite, an open platform for developing computer-use agents. It is built around three standardized abstractions — Lite.Gym for environments, Lite.Sample for supervised data, and one harness per model across eval, SFT and RL. 15+ benchmarks, 10+ datasets and 30k+ verifiable tasks are already plugged in.
Why CUA-Lite
Computer-use agents operate desktop, browser and mobile applications, and developing one takes environments with verifiable tasks, supervised data, and models, plus the harness that runs a model in an environment. Open-source efforts already exist for all three, but they remain scattered: each environment has its own interface and runtime; each supervised dataset its own format; each model its own action space, prompt format and rollout code.
Without a shared standard, connecting models to environments means rewriting the same code: an action mapping, a context window, and a rollout loop — so resources cannot be pooled and scaled, nor can evaluation, SFT and RL share one infrastructure.
DatasetsMind2WebGUIOdysseyBenchmarksOSWorldWebArena
CUA-Lite closes all three, in one place:
- Lite.Gym, the unified environment interface — 15+ benchmarks, plus optional VM-free desktop sandboxes holding 30k+ verifiable training tasks.
- Lite.Sample, the unified supervised data format — 10+ datasets and fresh rollouts, free on Hugging Face.
- Model harnesses shared across eval, SFT and RL — one per model, implemented for 14 model families.
Lite.Gym: Unified environment interface
Any agent plugs into any environment through one interface.
lite.gymlite.gym# lite.gym — every env behind one interface
class LiteBaseEnv:
metadata: LiteMetadata # platform · extra tools
async def reset() # first screenshot + task
async def step(actions) # execute, observe, reward
async def close()
# one step of the loop ↓
env = gym.make("osworld@<task_id>")
result = await env.step([
click(coordinate=[500, 400]), # normalized [0, 1000]
])
result.observation.screenshot_b64 # base64 png, no prefix
result.reward # float | None
result.terminated, result.truncated
exposes an environment as a Gym-style reset / step / close loop. The loop standardizes one observation format across environments — a screenshot, text, or both per tool call — and one GUI action space per platform, issued as tool calls: click and type on desktop and browser, tap and swipe on mobile, plus extra tools per environment such as bash in our desktop sandboxes. It wraps the environment's own runtime: a VM, a browser, or a mobile emulator.
EnvironmentsOSWorld · desktopWebArena · browserAndroidWorld · mobileWebGym · browseryours next
lite.gymlite.gym# lite.gym — every env behind one interface
class LiteBaseEnv:
metadata: LiteMetadata # platform · extra tools
async def reset() # first screenshot + task
async def step(actions) # execute, observe, reward
async def close()
# one step of the loop ↓
env = gym.make("osworld@<task_id>")
result = await env.step([
click(coordinate=[500, 400]), # normalized [0, 1000]
])
result.observation.screenshot_b64 # base64 png, no prefix
result.reward # float | None
result.terminated, result.truncated
AgentsGPTClaudeQwenGeminiyours next
Agents
screenshot click([632, 375])
$ rollout.py --model-idgpt-5.5--env-idosworld
CUA-Lite also provides optional, lightweight VM-free desktop sandboxes of its own: Docker containers that replicate OSWorld's desktop at much lower cost and hold 30k+ verifiable tasks for training. They need no /dev/kvm, the hardware virtualization VM-based benchmarks require, so they run on any host with Docker — and Lite.OSWorld (ours) runs OSWorld's own tasks and evaluators, unchanged.
OSWorld
Ubuntu.qcow2
QEMU · KVM
/dev/kvm
Lite.OSWorld · parallel rollouts
A desktop sealed in a VM — every task boots QEMU/KVM.
Beyond OSWorld: scalable training sandboxes
The VM-free container isn't just for OSWorld — it's the foundation for CUA-Lite's family of sandboxes. The same base already runs browser and desktop tasks, and real science desktops: GMAT flying spacecraft, PyMOL turning proteins.
Lite.OSWorld Lite.ScaleCUA Lite.CUAGym Lite.CUAWorld
369 benchmark tasks + 2k+ synthesized, across 10 desktop apps.
Looping rollout trajectories — click a tile for the full rollout.
Call for sandbox contributors. A sandbox only matters while people run it. Add yours to CUA-Lite, and every agent trains and benchmarks on it — now and later. One integration, and the whole field builds on it.
Lite.* environments ↗·Env guide ↗·Leaderboard ↓
Lite.Sample: Unified supervised data format
Convert a dataset once, and every agent can train on it.
LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
metadata: LiteMetadata # platform · task_type
images: list[Image] # screenshots
messages: list[LiteMessage] # user ⇄ assistant turns
# a grounding step — one action ↓
LiteSample(
LiteMetadata("desktop", "grounding.action"),
messages=[
user("Click Subscribe", img=0),
assistant(click([455, 215])),
])
# a use trajectory — 2 steps ↓
LiteSample(
LiteMetadata("browser", "use"),
messages=[
user("Find cua-lite on GitHub", img=0),
assistant(type("cua-lite")),
user(img=1), # result screenshot
assistant(click([320, 180]), terminate()),
])
is the one schema, shared across every env, agent, and task type, free on Hugging Face.
One shape for all of it, whatever it came from: messages whose tool calls are the interface's actions and whose tool results are its observations — from a GUI grounding label to a full rollout.
DatasetsMind2Web · browserGUIOdyssey · mobileScaleCUA · allOpenCUA · desktopyours
LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
metadata: LiteMetadata # platform · task_type
images: list[Image] # screenshots
messages: list[LiteMessage] # user ⇄ assistant turns
# a grounding step — one action ↓
LiteSample(
LiteMetadata("desktop", "grounding.action"),
messages=[
user("Click Subscribe", img=0),
assistant(click([455, 215])),
])
# a use trajectory — 2 steps ↓
LiteSample(
LiteMetadata("browser", "use"),
messages=[
user("Find cua-lite on GitHub", img=0),
assistant(type("cua-lite")),
user(img=1), # result screenshot
assistant(click([320, 180]), terminate()),
])
AgentsQwenUI-TARSMAI-UIFarayours
Agents
adapter adapter# the adapter — one per model family class BaseAgentAdapter: action_space: BaseActionSpace # its coordinates protocol: BaseProtocol # its history def render_step(sample, k, processed) # → the messages the model sees at turn k def unroll(sample) # → AgentSample, every turn
the same adapter serves rollout and export_sft —
the SFT prompt matches what the model saw live
$ export_sft --data-paths hf/Multimodal-Mind2Web--model-idQwen3-VL-8B-Instruct
**CUA-Lite ships an
adapteradapter# the adapter — one per model family
class BaseAgentAdapter:
action_space: BaseActionSpace # its coordinates
protocol: BaseProtocol # its history
def render_step(sample, k, processed)
# → the messages the model sees at turn k
def unroll(sample) # → AgentSample, every turn
# the same adapter serves rollout and export_sft —
# the SFT prompt matches what the model saw live
per model, packing a unified
LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
metadata: LiteMetadata # platform · task_type
images: list[Image] # screenshots
messages: list[LiteMessage] # user ⇄ assistant turns
# a grounding step — one action ↓
LiteSample(
LiteMetadata("desktop", "grounding.action"),
messages=[
user("Click Subscribe", img=0),
assistant(click([455, 215])),
])
# a use trajectory — 2 steps ↓
LiteSample(
LiteMetadata("browser", "use"),
messages=[
user("Find cua-lite on GitHub", img=0),
assistant(type("cua-lite")),
user(img=1), # result screenshot
assistant(click([320, 180]), terminate()),
])
into the exact training format each one needs.** The figure above shows one, with the building blocks to add your own.
10+ datasets are already on Hugging Face: existing CUA corpora — grounding · understanding · use — preprocessed into Lite.Sample, plus fresh rollouts from frontier CUAs. Browse the corpora and the rollouts; below, one of them, WebGym:
huggingface.co/datasets/cua-lite/WebGym
Rollouts
Lite.OSWorld WebGym Lite.ScaleCUA Lite.CUAWorld Lite.CUAGym
Corpora
Aguvis CAGUI GUI-360 GUIAct GUIOdyssey Multimodal-Mind2Web OpenCUA ScaleCUA UI-Genie-Agentopen ↗
Call for data contributors. Data only matters while models can train on it. Share yours with CUA-Lite, and every agent trains on it — even models that don't exist yet. One conversion, and the whole community trains on it.
Preprocessing guide ↗·Agent harnesses ↗·SFT guide ↗
Model harnesses: shared across eval, SFT & RL
Each model has one harness, the code that adapts it to the interface and the format. Through its harness a model runs in every integrated environment, and eval and RL consume the rollouts it produces; with the same harness, any stored
LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
metadata: LiteMetadata # platform · task_type
images: list[Image] # screenshots
messages: list[LiteMessage] # user ⇄ assistant turns
# a grounding step — one action ↓
LiteSample(
LiteMetadata("desktop", "grounding.action"),
messages=[
user("Click Subscribe", img=0),
assistant(click([455, 215])),
])
# a use trajectory — 2 steps ↓
LiteSample(
LiteMetadata("browser", "use"),
messages=[
user("Find cua-lite on GitHub", img=0),
assistant(type("cua-lite")),
user(img=1), # result screenshot
assistant(click([320, 180]), terminate()),
])
is rendered into the model's own training format for SFT — here, Qwen3.5's:
LiteSample{ }adapter Qwen
step 1 _instr · img1_ _act1_
step 2 _instr · img1_ _act1_ _img2_ _act2_
step 3 _instr · img1_ _act1_ _img2_ _act2_ _img3_ _act3_
step 4 _instr · img1_ _act1_ _img2_ _act2_ _img3_ _act3_ _img4_ _act4_ 1 forward · loss ×4
step 5 _hist ×4_ _img5_ _act5_
step 6 _hist ×4_ _img5_ _act5_ _img6_ _act6_ 1 forward · loss ×2
prompt response · loss collapsed history
one trajectory, 6 steps — as Qwen3.5 sees it
Eval, any agent on any benchmark
Set --model-id for the agent and --env-id for the benchmark:
evaluate.sh
$ python scripts/rollout.py \
--model-id gpt-5.5 gpt-5.5 gpt-5.4 claude-opus-4-7 claude-opus-4-6 claude-sonnet-4-6 Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview\
--env-id osworld Grounding screenspot_pro osworld_g Desktop osworld lite.osworld osworld_2 cua.bench Browser webgym webharbor.webvoyager online_mind2web browsergym.miniwob browsergym.webarena browsergym.visualwebarena Mobile androidworld androidlab mobileworld mobilegym\
--splits eval \
--config-path scripts/configs/gpt/default/osworld.yaml
15+ benchmarks are already integrated — ours is the VM-free runtime, the interface and the integration, not the task suites — and the VM-free desktop sandboxes are an addition, not a replacement: the original OSWorld VM sits right beside Lite.OSWorld, and the mobile benchmarks still need a VM or an emulator:
Grounding 2 Desktop 4 Browser 6 Mobile 4
OSWorld-G desktop UI groundingScreenSpot-Pro high-res professional apps
OSWorld real Ubuntu desktop appsLite.OSWorld lightweight OSWorld tasksOSWorld-2 capability-graded · partial creditCUABench computer-use · KiCad · workflows
WebGym web navigationWebVoyager live production websitesOnline-Mind2Web 300 live web tasksMiniWoB mini web interactionsWebArena self-hosted shop · forum · GitLabVisualWebArena visual web tasks
AndroidWorld real Android appsAndroidLab 9 offline Android appsMobileWorld 161 mobile tasksMobileGym mobile navigation
Leaderboard OSWorld compositionREADME↗
1gpt-5.572.3%
2Qwen/Qwen3.5-27B46.2%
3meituan/EvoCUA-8B-2026010542.8%
4Qwen/Qwen3.5-9B38.1%
5Qwen/Qwen3-VL-32B-Instruct35.8%
6Qwen/Qwen3-VL-8B-Instruct34.2%
7Qwen/Qwen3.5-4B30.3%
8ByteDance-Seed/UI-TARS-1.5-7B29.4%
9Qwen/Qwen3-VL-4B-Instruct29.1%
10Qwen/Qwen3-VL-2B-Instruct20.5%
11OpenGVLab/ScaleCUA-7B15.0%
12ByteDance-Seed/UI-TARS-7B-DPO14.7%
13Qwen/Qwen3.5-2B10.7%
13 agents 325 tasks success rate 2026-07-18 · run_0.json @ 65ea6496
Composition
OSWorld
OSWorld — 10 real Ubuntu desktop apps.
upstream 369
excluded infeasible / broken evaluator−48
scored 321
SFT & RL, any open agent
SFT on CUA-Lite's corpora, then reinforce in its envs — GRPO and beyond, on the Slime trainer. Train any open agent on any data and any env:
SFT RL
Lite.Sample, adapted to each model — pick a dataset and a student:
Dataset:ScaleCUA Rollouts Lite.OSWorld WebGym Lite.ScaleCUA Lite.CUAWorld Lite.CUAGym Corpora Aguvis CAGUI GUI-360 GUIAct GUIOdyssey Multimodal-Mind2Web OpenCUA ScaleCUA UI-Genie-Agent
Model:Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview
run_sft.sh
--- host ---
1 · download the corpus
$ python -m lite.data.hf.download ScaleCUA\
--out .data/hf/cua-lite/ScaleCUA
2 · export a model-ready SFT parquet
$ python -m lite.train.export.export_sft \
--model-id Qwen/Qwen3-VL-8B-Instruct\
--config scripts/configs/qwen3_vl/recipes/sft/default.yaml \
--data-paths .data/hf/cua-lite/ScaleCUA\
--image-root .data/hf \
-o .data/sft/qwen3_vl/scalecua.parquet
--- Slime container (see docs/slime.md) ---
3 · supervised fine-tune
$MODEL_ID=Qwen/Qwen3-VL-8B-Instruct\
PROMPT_DATA=.data/sft/qwen3_vl/scalecua.parquet \
bash scripts/train/run_sft.sh
Rollouts scored in the env drive GRPO updates — pick a model and env:
run_grpo.sh
$MODEL_ID=Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview\
ENV_ID=lite.osworld Grounding screenspot_pro Desktop lite.osworld Browser webgym Mobile androidworld mobilegym\
CONFIG_PATH=scripts/configs/qwen3_vl/default/lite.osworld.yaml \
bash scripts/train/run_grpo.sh
Bring a dataset, an env, or an agent — each one compounds. Or just tell us what's missing — GitHub · Hugging Face · Email.
来源:Berkeley RDI:Blog(AI 安全与评测) · rdi.berkeley.edu