Hugging Face 发布 ScreenEnv:用 Docker 容器部署桌面 GUI Agent
ScreenEnv: Deploy your full stack Desktop Agent
Hugging Face 发布开源 Python 库 ScreenEnv,可在 Docker 容器中创建隔离的 Ubuntu 桌面环境,用于测试和部署 GUI Agent(Computer Use)。
官方介绍了 Docker 化桌面环境的部署速度、双集成模式和 smolagents 用法,方便开发者按自己架构搭建桌面 Agent。
TL;DR:ScreenEnv 是一个强大的 Python 库,让你可以在 Docker 容器中创建隔离的 Ubuntu 桌面环境,用于测试和部署 GUI Agent(即 Computer Use agent)。借助对 Model Context Protocol(MCP)的内置支持,部署能够看到、点击并与真实应用交互的桌面 agent 从未如此简单。
什么是 ScreenEnv?
假设你需要自动化桌面任务、测试 GUI 应用,或者构建一个能够与软件交互的 AI agent。过去这需要复杂的 VM 配置和脆弱的自动化框架。
ScreenEnv 通过提供一个运行在 Docker 容器中的沙盒化桌面环境改变了这一点。可以把它看作一个完整的虚拟桌面会话,你的代码可以完全控制它——不仅仅是点击按钮和输入文本,还能管理整个桌面体验,包括启动应用、整理窗口、处理文件、执行终端命令以及录制整个会话。
为什么选择 ScreenEnv?
- 🖥️ 完整桌面控制:完整的鼠标和键盘自动化、窗口管理、应用启动、文件操作、终端访问以及屏幕录制
- 🤖 双重集成模式:同时支持面向 AI 系统的 Model Context Protocol(MCP)和直接的 Sandbox API——可适配任何 agent 或后端逻辑
- 🐳 Docker 原生:无需复杂的 VM 配置——只需 Docker。环境隔离、可复现,并可在不到 10 秒内部署到任何地方。支持 AMD64 和 ARM64 架构。
🎯 一行命令完成设置
from screenenv import Sandbox
sandbox = Sandbox() # That's it!
两种集成方式
ScreenEnv 提供两种互补的方式来与你的 agent 和后端系统集成,让你可以灵活选择最适合自己架构的方式:
方式 1:直接使用 Sandbox API
非常适合自定义 agent 框架、现有后端,或者当你需要细粒度控制时:
from screenenv import Sandbox
# Direct programmatic control
sandbox = Sandbox(headless=False)
sandbox.launch("xfce4-terminal")
sandbox.write("echo 'Custom agent logic'")
screenshot = sandbox.screenshot()
image = Image.open(BytesIO(screenshot_bytes))
...
sandbox.close()
# If close() isn’t called, you might need to shut down the container yourself.
方式 2:MCP Server 集成
非常适合支持 Model Context Protocol 的 AI 系统:
from screenenv import MCPRemoteServer
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client
# Start MCP server for AI integration
server = MCPRemoteServer(headless=False)
print(f"MCP Server URL: {server.server_url}")
# AI agents can now connect and control the desktop
async def mcp_session():
async with streamablehttp_client(server.server_url) as streams:
async with ClientSession(*streams) as session:
await session.initialize()
print(await session.list_tools())
response = await session.call_tool("screenshot", {})
image_bytes = base64.b64decode(response.content[0].data)
image = Image.open(BytesIO(image_bytes))
server.close()
# If close() isn’t called, you might need to shut down the container yourself.
这种双重方式意味着 ScreenEnv 能够适配你现有的基础设施,而不是强迫你改变 agent 架构。
✨ 使用 screenenv 和 smolagents 创建桌面 Agent
screenenv 原生支持 smolagents,让你可以轻松构建自己的自定义桌面 Agent 以实现自动化。以下是如何在短短几步内创建你自己的 AI 驱动桌面 Agent:
1. 选择你的模型
选择你希望用来驱动 agent 的后端 VLM。
import os
from smolagents import OpenAIServerModel
model = OpenAIServerModel(
model_id="gpt-4.1",
api_key=os.getenv("OPENAI_API_KEY"),
)
# Inference Endpoints
from smolagents import HfApiModel
model = HfApiModel(
model_id="Qwen/Qwen2.5-VL-7B-Instruct",
token=os.getenv("HF_TOKEN"),
provider="nebius",
)
# Transformer models
from smolagents import TransformersModel
model = TransformersModel(
model_id="Qwen/Qwen2.5-VL-7B-Instruct",
device_map="auto",
torch_dtype="auto",
trust_remote_code=True,
)
# Other providers
from smolagents import LiteLLMModel
model = LiteLLMModel(model_id="anthropic/claude-sonnet-4-20250514")
# see smolagents to get the list of available model connectors
2. 定义你的自定义桌面 Agent
继承 DesktopAgentBase 并实现 _setup_desktop_tools 方法,以构建你自己的动作空间!
from screenenv import DesktopAgentBase, Sandbox
from smolagents import Model, Tool, tool
from smolagents.monitoring import LogLevel
from typing import List
class CustomDesktopAgent(DesktopAgentBase):
"""Agent for desktop automation"""
def __init__(
self,
model: Model,
data_dir: str,
desktop: Sandbox,
tools: List[Tool] | None = None,
max_steps: int = 200,
verbosity_level: LogLevel = LogLevel.INFO,
planning_interval: int | None = None,
use_v1_prompt: bool = False,
**kwargs,
):
super().__init__(
model=model,
data_dir=data_dir,
desktop=desktop,
tools=tools,
max_steps=max_steps,
verbosity_level=verbosity_level,
planning_interval=planning_interval,
use_v1_prompt=use_v1_prompt,
**kwargs,
)
# OPTIONAL: Add a custom prompt template - see src/screenenv/desktop_agent/desktop_agent_base.py for more details about the default prompt template
# self.prompt_templates["system_prompt"] = CUSTOM_PROMPT_TEMPLATE.replace(
# "<<resolution_x>>", str(self.width)
# ).replace("<<resolution_y>>", str(self.height))
# Important: Adjust the prompt based on your action space to improve results.
def _setup_desktop_tools(self) -> None:
"""Define your custom tools here."""
@tool
def click(x: int, y: int) -> str:
"""
Clicks at the specified coordinates.
Args:
x: The x-coordinate of the click
y: The y-coordinate of the click
"""
self.desktop.left_click(x, y)
# self.click_coordinates = (x, y) to add the click coordinate to the observation screenshot
return f"Clicked at ({x}, {y})"
self.tools["click"] = click
@tool
def write(text: str) -> str:
"""
Types the specified text at the current cursor position.
Args:
text: The text to type
"""
self.desktop.write(text, delay_in_ms=10)
return f"Typed text: '{text}'"
self.tools["write"] = write
@tool
def press(key: str) -> str:
"""
Presses a keyboard key or combination of keys
Args:
key: The key to press (e.g. "enter", "space", "backspace", etc.) or a multiple keys string to press, for example "ctrl+a" or "ctrl+shift+a".
"""
self.desktop.press(key)
return f"Pressed key: {key}"
self.tools["press"] = press
@tool
def open(file_or_url: str) -> str:
"""
Directly opens a browser with the specified url or opens a file with the default application.
Args:
file_or_url: The URL or file to open
"""
self.desktop.open(file_or_url)
# Give it time to load
self.logger.log(f"Opening: {file_or_url}")
return f"Opened: {file_or_url}"
@tool
def launch_app(app_name: str) -> str:
"""
Launches the specified application.
Args:
app_name: The name of the application to launch
"""
self.desktop.launch(app_name)
return f"Launched application: {app_name}"
self.tools["launch_app"] = launch_app
... # Continue implementing your own action space.
3. 在桌面任务上运行 Agent
from screenenv import Sandbox
# Define your sandbox environment
sandbox = Sandbox(headless=False, resolution=(1920, 1080))
# Create your agent
agent = CustomDesktopAgent(
model=model,
data_dir="data",
desktop=sandbox,
)
# Run a task
task = "Open LibreOffice, write a report of approximately 300 words on the topic ‘AI Agent Workflow in 2025’, and save the document."
result = agent.run(task)
print(f"📄 Result: {result}")
sandbox.close()
如果你遇到 access denied docker 错误,可以尝试使用
sudo -E python -m test.py运行 agent,或者将你的用户添加到docker组。
💡 如需完整实现,请查看 GitHub 上的这个 CustomDesktopAgent 源码。
立即开始
# Install ScreenEnv
pip install screenenv
# Try the examples
git clone git@github.com:huggingface/screenenv.git
cd screenenv
python -m examples.desktop_agent
# use 'sudo -E python -m examples.desktop_agent` if you're not in 'docker' group
下一步是什么?
ScreenEnv 的目标是超越 Linux,支持 Android、macOS 和 Windows,从而实现真正的跨平台 GUI 自动化。这将使开发者和研究人员能够以最少的设置构建可跨环境泛化的 agent。
这些进步为创建可复现的沙盒环境铺平了道路,非常适合用于基准测试和评估。
来源:Hugging Face:Blog · huggingface.co