跳到正文
北京时间
原文
Berkeley RDI:Blog(AI 安全与评测)·· 23 天前精选AI 评分61

Berkeley RDI 发布开源平台 CUA-Lite,面向计算机使用智能体

CUA-Lite: An Open Platform for Computer-Use Agents

AI 导读

Berkeley RDI 推出开源平台 CUA-Lite,为计算机使用智能体提供三个标准化抽象。其环境接口下运行 15+ 个覆盖桌面、浏览器和移动端的基准,并提供含 30k+ 可验证任务的免 VM 桌面沙箱;统一监督数据格式已转换 10+ 个公开 CUA 数据集;每模型一个 harness 在评测、SFT 和 RL 间共享,支持 14 个模型族。

推荐理由

原文说明了平台的三项标准化抽象及覆盖规模,读者可以据此评估它在整合分散的计算机使用智能体资源上的可用性。

正文 · 原文

Introducing CUA-Lite, an open platform for developing computer-use agents. It is built around three standardized abstractions — Lite.Gym for environments, Lite.Sample for supervised data, and one harness per model across eval, SFT and RL. 15+ benchmarks, 10+ datasets and 30k+ verifiable tasks are already plugged in.

Code ↗·Project site ↗

Why CUA-Lite

Computer-use agents operate desktop, browser and mobile applications, and developing one takes environments with verifiable tasks, supervised data, and models, plus the harness that runs a model in an environment. Open-source efforts already exist for all three, but they remain scattered: each environment has its own interface and runtime; each supervised dataset its own format; each model its own action space, prompt format and rollout code.

Without a shared standard, connecting models to environments means rewriting the same code: an action mapping, a context window, and a rollout loop — so resources cannot be pooled and scaled, nor can evaluation, SFT and RL share one infrastructure.

DatasetsMind2WebGUIOdysseyBenchmarksOSWorldWebArena

AgentsGPTClaudeQwenGemini

CUA-Lite closes all three, in one place:

  • Lite.Gym, the unified environment interface — 15+ benchmarks, plus optional VM-free desktop sandboxes holding 30k+ verifiable training tasks.
  • Lite.Sample, the unified supervised data format — 10+ datasets and fresh rollouts, free on Hugging Face.
  • Model harnesses shared across eval, SFT and RL — one per model, implemented for 14 model families.

Lite.Gym: Unified environment interface

Any agent plugs into any environment through one interface.

lite.gymlite.gym# lite.gym — every env behind one interface
class LiteBaseEnv:
    metadata: LiteMetadata     # platform · extra tools
    async def reset()        # first screenshot + task
    async def step(actions)  # execute, observe, reward
    async def close()

# one step of the loop ↓
env = gym.make("osworld@<task_id>")
result = await env.step([
  click(coordinate=[500, 400]),  # normalized [0, 1000]
])
result.observation.screenshot_b64   # base64 png, no prefix
result.reward                       # float | None
result.terminated, result.truncated

exposes an environment as a Gym-style reset / step / close loop. The loop standardizes one observation format across environments — a screenshot, text, or both per tool call — and one GUI action space per platform, issued as tool calls: click and type on desktop and browser, tap and swipe on mobile, plus extra tools per environment such as bash in our desktop sandboxes. It wraps the environment's own runtime: a VM, a browser, or a mobile emulator.

EnvironmentsOSWorld · desktopWebArena · browserAndroidWorld · mobileWebGym · browseryours next

lite.gymlite.gym# lite.gym — every env behind one interface
class LiteBaseEnv:
    metadata: LiteMetadata     # platform · extra tools
    async def reset()        # first screenshot + task
    async def step(actions)  # execute, observe, reward
    async def close()

# one step of the loop ↓
env = gym.make("osworld@<task_id>")
result = await env.step([
  click(coordinate=[500, 400]),  # normalized [0, 1000]
])
result.observation.screenshot_b64   # base64 png, no prefix
result.reward                       # float | None
result.terminated, result.truncated

AgentsGPTClaudeQwenGeminiyours next

Agents

screenshot click([632, 375])

$ rollout.py --model-idgpt-5.5--env-idosworld

CUA-Lite also provides optional, lightweight VM-free desktop sandboxes of its own: Docker containers that replicate OSWorld's desktop at much lower cost and hold 30k+ verifiable tasks for training. They need no /dev/kvm, the hardware virtualization VM-based benchmarks require, so they run on any host with Docker — and Lite.OSWorld (ours) runs OSWorld's own tasks and evaluators, unchanged.

OSWorld

Ubuntu.qcow2

QEMU · KVM

/dev/kvm

Lite.OSWorld · parallel rollouts

A desktop sealed in a VM — every task boots QEMU/KVM.

Beyond OSWorld: scalable training sandboxes

The VM-free container isn't just for OSWorld — it's the foundation for CUA-Lite's family of sandboxes. The same base already runs browser and desktop tasks, and real science desktops: GMAT flying spacecraft, PyMOL turning proteins.

Lite.OSWorld Lite.ScaleCUA Lite.CUAGym Lite.CUAWorld

369 benchmark tasks + 2k+ synthesized, across 10 desktop apps.

Looping rollout trajectories — click a tile for the full rollout.

Call for sandbox contributors. A sandbox only matters while people run it. Add yours to CUA-Lite, and every agent trains and benchmarks on it — now and later. One integration, and the whole field builds on it.

Lite.* environments ↗·Env guide ↗·Leaderboard ↓

Lite.Sample: Unified supervised data format

Convert a dataset once, and every agent can train on it.

LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
    metadata: LiteMetadata       # platform · task_type
    images:   list[Image]        # screenshots
    messages: list[LiteMessage]  # user ⇄ assistant turns

# a grounding step — one action ↓
LiteSample(
  LiteMetadata("desktop", "grounding.action"),
  messages=[
    user("Click Subscribe", img=0),
    assistant(click([455, 215])),
  ])

# a use trajectory — 2 steps ↓
LiteSample(
  LiteMetadata("browser", "use"),
  messages=[
    user("Find cua-lite on GitHub", img=0),
    assistant(type("cua-lite")),
    user(img=1),                  # result screenshot
    assistant(click([320, 180]), terminate()),
  ])

is the one schema, shared across every env, agent, and task type, free on Hugging Face.

One shape for all of it, whatever it came from: messages whose tool calls are the interface's actions and whose tool results are its observations — from a GUI grounding label to a full rollout.

DatasetsMind2Web · browserGUIOdyssey · mobileScaleCUA · allOpenCUA · desktopyours

LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
    metadata: LiteMetadata       # platform · task_type
    images:   list[Image]        # screenshots
    messages: list[LiteMessage]  # user ⇄ assistant turns

# a grounding step — one action ↓
LiteSample(
  LiteMetadata("desktop", "grounding.action"),
  messages=[
    user("Click Subscribe", img=0),
    assistant(click([455, 215])),
  ])

# a use trajectory — 2 steps ↓
LiteSample(
  LiteMetadata("browser", "use"),
  messages=[
    user("Find cua-lite on GitHub", img=0),
    assistant(type("cua-lite")),
    user(img=1),                  # result screenshot
    assistant(click([320, 180]), terminate()),
  ])

AgentsQwenUI-TARSMAI-UIFarayours

Agents

adapter adapter# the adapter — one per model family class BaseAgentAdapter: action_space: BaseActionSpace # its coordinates protocol: BaseProtocol # its history def render_step(sample, k, processed) # → the messages the model sees at turn k def unroll(sample) # → AgentSample, every turn

the same adapter serves rollout and export_sft —

the SFT prompt matches what the model saw live

$ export_sft --data-paths hf/Multimodal-Mind2Web--model-idQwen3-VL-8B-Instruct

**CUA-Lite ships an

adapteradapter# the adapter — one per model family
class BaseAgentAdapter:
    action_space: BaseActionSpace  # its coordinates
    protocol: BaseProtocol         # its history
    def render_step(sample, k, processed)
        # → the messages the model sees at turn k
    def unroll(sample)  # → AgentSample, every turn

# the same adapter serves rollout and export_sft —
# the SFT prompt matches what the model saw live

per model, packing a unified

LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
    metadata: LiteMetadata       # platform · task_type
    images:   list[Image]        # screenshots
    messages: list[LiteMessage]  # user ⇄ assistant turns

# a grounding step — one action ↓
LiteSample(
  LiteMetadata("desktop", "grounding.action"),
  messages=[
    user("Click Subscribe", img=0),
    assistant(click([455, 215])),
  ])

# a use trajectory — 2 steps ↓
LiteSample(
  LiteMetadata("browser", "use"),
  messages=[
    user("Find cua-lite on GitHub", img=0),
    assistant(type("cua-lite")),
    user(img=1),                  # result screenshot
    assistant(click([320, 180]), terminate()),
  ])

into the exact training format each one needs.** The figure above shows one, with the building blocks to add your own.

10+ datasets are already on Hugging Face: existing CUA corpora — grounding · understanding · use — preprocessed into Lite.Sample, plus fresh rollouts from frontier CUAs. Browse the corpora and the rollouts; below, one of them, WebGym:

huggingface.co/datasets/cua-lite/WebGym

Rollouts

Lite.OSWorld WebGym Lite.ScaleCUA Lite.CUAWorld Lite.CUAGym

Corpora

Aguvis CAGUI GUI-360 GUIAct GUIOdyssey Multimodal-Mind2Web OpenCUA ScaleCUA UI-Genie-Agentopen ↗

Call for data contributors. Data only matters while models can train on it. Share yours with CUA-Lite, and every agent trains on it — even models that don't exist yet. One conversion, and the whole community trains on it.

Preprocessing guide ↗·Agent harnesses ↗·SFT guide ↗

Model harnesses: shared across eval, SFT & RL

Each model has one harness, the code that adapts it to the interface and the format. Through its harness a model runs in every integrated environment, and eval and RL consume the rollouts it produces; with the same harness, any stored

LiteSampleLiteSample# LiteSample — one schema for any data
@dataclass
class LiteSample:
    metadata: LiteMetadata       # platform · task_type
    images:   list[Image]        # screenshots
    messages: list[LiteMessage]  # user ⇄ assistant turns

# a grounding step — one action ↓
LiteSample(
  LiteMetadata("desktop", "grounding.action"),
  messages=[
    user("Click Subscribe", img=0),
    assistant(click([455, 215])),
  ])

# a use trajectory — 2 steps ↓
LiteSample(
  LiteMetadata("browser", "use"),
  messages=[
    user("Find cua-lite on GitHub", img=0),
    assistant(type("cua-lite")),
    user(img=1),                  # result screenshot
    assistant(click([320, 180]), terminate()),
  ])

is rendered into the model's own training format for SFT — here, Qwen3.5's:

LiteSample{ }adapter Qwen

step 1 _instr · img1_ _act1_

step 2 _instr · img1_ _act1_ _img2_ _act2_

step 3 _instr · img1_ _act1_ _img2_ _act2_ _img3_ _act3_

step 4 _instr · img1_ _act1_ _img2_ _act2_ _img3_ _act3_ _img4_ _act4_ 1 forward · loss ×4

step 5 _hist ×4_ _img5_ _act5_

step 6 _hist ×4_ _img5_ _act5_ _img6_ _act6_ 1 forward · loss ×2

prompt response · loss collapsed history

one trajectory, 6 steps — as Qwen3.5 sees it

Eval, any agent on any benchmark

Set --model-id for the agent and --env-id for the benchmark:

evaluate.sh

$ python scripts/rollout.py \

--model-id gpt-5.5 gpt-5.5 gpt-5.4 claude-opus-4-7 claude-opus-4-6 claude-sonnet-4-6 Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview\

--env-id osworld Grounding screenspot_pro osworld_g Desktop osworld lite.osworld osworld_2 cua.bench Browser webgym webharbor.webvoyager online_mind2web browsergym.miniwob browsergym.webarena browsergym.visualwebarena Mobile androidworld androidlab mobileworld mobilegym\

--splits eval \

--config-path scripts/configs/gpt/default/osworld.yaml

15+ benchmarks are already integrated — ours is the VM-free runtime, the interface and the integration, not the task suites — and the VM-free desktop sandboxes are an addition, not a replacement: the original OSWorld VM sits right beside Lite.OSWorld, and the mobile benchmarks still need a VM or an emulator:

Grounding 2 Desktop 4 Browser 6 Mobile 4

OSWorld-G desktop UI groundingScreenSpot-Pro high-res professional apps

OSWorld real Ubuntu desktop appsLite.OSWorld lightweight OSWorld tasksOSWorld-2 capability-graded · partial creditCUABench computer-use · KiCad · workflows

WebGym web navigationWebVoyager live production websitesOnline-Mind2Web 300 live web tasksMiniWoB mini web interactionsWebArena self-hosted shop · forum · GitLabVisualWebArena visual web tasks

AndroidWorld real Android appsAndroidLab 9 offline Android appsMobileWorld 161 mobile tasksMobileGym mobile navigation

Leaderboard OSWorld compositionREADME↗

1gpt-5.572.3%

2Qwen/Qwen3.5-27B46.2%

3meituan/EvoCUA-8B-2026010542.8%

4Qwen/Qwen3.5-9B38.1%

5Qwen/Qwen3-VL-32B-Instruct35.8%

6Qwen/Qwen3-VL-8B-Instruct34.2%

7Qwen/Qwen3.5-4B30.3%

8ByteDance-Seed/UI-TARS-1.5-7B29.4%

9Qwen/Qwen3-VL-4B-Instruct29.1%

10Qwen/Qwen3-VL-2B-Instruct20.5%

11OpenGVLab/ScaleCUA-7B15.0%

12ByteDance-Seed/UI-TARS-7B-DPO14.7%

13Qwen/Qwen3.5-2B10.7%

13 agents 325 tasks success rate 2026-07-18 · run_0.json @ 65ea6496

Composition

OSWorld

OSWorld — 10 real Ubuntu desktop apps.

upstream 369

excluded infeasible / broken evaluator−48

scored 321

SFT & RL, any open agent

SFT on CUA-Lite's corpora, then reinforce in its envs — GRPO and beyond, on the Slime trainer. Train any open agent on any data and any env:

SFT RL

Lite.Sample, adapted to each model — pick a dataset and a student:

Dataset:ScaleCUA Rollouts Lite.OSWorld WebGym Lite.ScaleCUA Lite.CUAWorld Lite.CUAGym Corpora Aguvis CAGUI GUI-360 GUIAct GUIOdyssey Multimodal-Mind2Web OpenCUA ScaleCUA UI-Genie-Agent

Model:Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview

run_sft.sh

--- host ---

1 · download the corpus

$ python -m lite.data.hf.download ScaleCUA\

--out .data/hf/cua-lite/ScaleCUA

2 · export a model-ready SFT parquet

$ python -m lite.train.export.export_sft \

--model-id Qwen/Qwen3-VL-8B-Instruct\

--config scripts/configs/qwen3_vl/recipes/sft/default.yaml \

--data-paths .data/hf/cua-lite/ScaleCUA\

--image-root .data/hf \

-o .data/sft/qwen3_vl/scalecua.parquet

--- Slime container (see docs/slime.md) ---

3 · supervised fine-tune

$MODEL_ID=Qwen/Qwen3-VL-8B-Instruct\

PROMPT_DATA=.data/sft/qwen3_vl/scalecua.parquet \

bash scripts/train/run_sft.sh

Rollouts scored in the env drive GRPO updates — pick a model and env:

run_grpo.sh

$MODEL_ID=Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-4B-Instruct Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3.5-4B Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B Qwen/Qwen3.5-2B Qwen/Qwen2.5-VL-7B-Instruct Qwen/Qwen2.5-VL-3B-Instruct microsoft/Fara-7B ByteDance-Seed/UI-TARS-1.5-7B ByteDance-Seed/UI-TARS-7B-DPO OpenGVLab/ScaleCUA-7B xlangai/OpenCUA-7B meituan/EvoCUA-8B-20260105 Tongyi-MAI/MAI-UI-8B Tongyi-MAI/MAI-UI-2B MarsXL/UI-Voyager stepfun-ai/GELab-Zero-4B-preview\

ENV_ID=lite.osworld Grounding screenspot_pro Desktop lite.osworld Browser webgym Mobile androidworld mobilegym\

CONFIG_PATH=scripts/configs/qwen3_vl/default/lite.osworld.yaml \

bash scripts/train/run_grpo.sh

Bring a dataset, an env, or an agent — each one compounds. Or just tell us what's missing — GitHub · Hugging Face · Email.

来源:Berkeley RDI:Blog(AI 安全与评测) · rdi.berkeley.edu