跳到正文
北京时间
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-06-08精选AI 评分73

OmniGameArena:面向VLM游戏智能体的统一UE5基准与改善动态

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

AI 导读

OmniGameArena是一个基于十二个Unreal Engine 5新构建游戏的实时基准,涵盖单人(7个)、PvP(3个)和合作(2个)模式,提供统一动作接口。除冷启动排行榜分数外,还引入Improvement Dynamics Curve (IDC),一种智能体反射评估机制:通过工具调用反射大语言模型自动优化技能提示词,追踪多轮反射中的分数变化以及习得技能在任务变体上的泛化表现。论文报告了12个VLM智能体在冷启动排行榜上的表现,以及4个顶级智能体在IDC下的指标。

推荐理由

在 UE5 里直接测 agent 的自我改进,这个思路让游戏 benchmark 从一次性的刷榜变成动态成长观测,对做多模态 agent 的团队是个新标尺。

正文 · 原文

Mingxian Lin

Shengju Qian

Yuqi Liu

Yi-Hua Huang

Yiyu Wang

Wei Huang

Yitang Li

Fan Zhang

Zeyu Hu

Lingting Zhu

Xin Wang

Xiaojuan Qi

The University of Hong Kong,

LIGHTSPEED,

Project Leader Corresponding Author

Abstract

Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-agent Solo play, and lack unified protocols for evaluating heterogeneous agent classes (commercial VLMs, open-weight VLMs, and specialized game policies) on the same footing. We address these gaps with OmniGameArena, a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo (7), PvP (3), and Coop (2) with unified action interfaces, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds. Beyond cold-start leaderboard scores, IDC exposes two additional observables for each (agent, game) pair: how the score evolves across reflection rounds, and how the learned skill behaves on held-out task variants. We report these observables for twelve VLM agents on the cold-start leaderboard and four top agents under IDC.

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

Mingxian Lin1, Shengju Qian2, ‡, Yuqi Liu3, Yi-Hua Huang1, Yiyu Wang2, Wei Huang1, Yitang Li4, Fan Zhang3, Zeyu Hu2, Lingting Zhu2, Xin Wang2, Xiaojuan Qi1, † 1The University of Hong Kong, 2LIGHTSPEED, 3The Chinese University of Hong Kong, 4Tsinghua University  Project Leader    Corresponding Author Project Page: https://mxlin043.github.io/OmniGameArena/

Refer to caption
Figure 1: OmniGameArena at a glance. Twelve newly built UE5 games span Solo (7), PvP (3), and Coop (2) regimes (top). Heterogeneous agents (commercial VLMs, open-weight VLMs, keyboard-mouse policies, and gamepad policies) connect to the same real-time UE5 environment through documented adapters (middle). Evaluation reports the cold-start leaderboard and the Improvement Dynamics Curve (IDC) under multi-round reflection (bottom).

1 Introduction

Foundation models are increasingly evaluated by how they act, not only by what they answer, and games are a natural stress test for this shift (Wang et al., 2023; Tan et al., 2024; Paglieri et al., 2024): an agent must read a changing visual scene, choose actions under time pressure, plan across delayed rewards, and adapt when the environment resists. Game benchmarks now span text-only worlds, 2D grid suites, and 3D open environments built on existing commercial titles, and have driven rapid progress in vision-language game agents (Tan et al., 2025; Magne et al., 2026; Wang et al., 2025b).

Yet current benchmarks rarely measure two properties that matter for deploying these agents. Most report a single first-attempt score per (agent, game) pair, leaving invisible the trajectory by which an agent improves under repeated interaction with the same task. They also lean heavily toward single-agent Solo play, while adversarial (PvP) and cooperative (Coop) regimes remain underrepresented even though they probe distinct capabilities such as opponent modeling, role assignment, and recovery from a teammate’s mistakes. Whether an agent can adapt under repeated reflection, and whether it can do so in adversarial or cooperative settings, therefore remains largely unmeasured.

We address both with OmniGameArena, a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo, PvP and Coop, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness built on top of it. The twelve games are authored for this benchmark rather than reused from public titles, lowering the risk of pre-training leakage, and share unified action interfaces (keyboard-mouse, gamepad) so that commercial VLMs, open-weight VLMs, and specialized game policies can all be evaluated under matched environment conditions. The IDC harness runs each (agent, game) instance for multiple rounds: the agent plays episodes under a current skill prompt, after which a reflector LLM inspects the trajectories through tool-use, deciding on its own what to read and when to stop, before refining the skill for the next round. We report both the per-round score sequence (the IDC of that instance) and a transfer score on held-out task variants.

Across twelve agents on the cold-start leaderboard, no single VLM dominates, and commercial agents hold a wide gap over open-weight VLMs and specialized policies. Among the four top agents that we run through IDC, all four improve over their cold-start baseline through reflection, yet peak performance is typically reached mid-curve rather than at the final round. Most notably, origin-task improvement and held-out variant transfer can diverge in our experiments; this divergence is hidden by single-round leaderboard scores and is a central observable IDC exposes.

To summarize, our contributions are threefold: (i) OmniGameArena, a twelve-game UE5 benchmark spanning Solo, PvP, and Coop with unified action interfaces and game instances built specifically for this benchmark; (ii) the IDC harness, an agentic-reflection framework whose autonomous tool-use reflector refines a bounded skill prompt across rounds, with persistent memory and best-skill rollback; and (iii) an empirical study across twelve agents showing that leadership rotates across games and that origin-task gain does not by itself predict held-out variant transfer.

Refer to caption
Figure 2: Radar charts of the 12 OmniGameArena games across seven capability dimensions. The abbreviations are: VP = Visual Perception, SN = Spatial Navigation, RT = Reaction, MEM = Memory, PLN = Planning, ADV = Adversarial, and COOP = Cooperation. Each dimension is scored from 0 to 3.

2 Related Work

Benchmarks in game environments. Interactive games have served as AI testbeds since the rise of reinforcement learning and now anchor evaluations of Large Language Models (LLMs) and Vision-Language Models (VLMs). Early LLM evaluations were text-only (Huang et al., 2024; Wu et al., 2023; Hu et al., 2024), effective for logical reasoning but lacking visual grounding. 2D-grid suites such as BALROG (Paglieri et al., 2024) and LVLM-Playground (Wang et al., 2025a) added spatial and multimodal demands, and more recent suites built on Minecraft or general visual benchmarks (V-MAGE (Zheng et al., 2025), Cradle (Tan et al., 2024), VideoGameBench (Zhang et al., 2025)) extended evaluation into 3D open worlds with long-horizon planning from pixels. Beyond game playing, related benchmarks probe embodied reasoning and action in complex 3D environments (Lin et al., 2025; Zhu et al., 2026) and unified reasoning across video-generation models (Luo et al., 2025), but neither targets multi-regime, real-time game interaction. Two limitations of these benchmarks motivate our work: most reuse existing commercial titles, leaving them exposed to pre-training contamination; and few cover Solo, PvP, and Coop regimes in a single real-time environment. OmniGameArena addresses both with twelve newly built UE5 games that span all three interaction regimes.

Game-playing LLM and VLM agents. Early LLM game agents operated in text-only environments (Hausknecht et al., 2020; Tsai et al., 2023) and 2D grid worlds (Feng et al., 2023; Küttler et al., 2020). Voyager (Wang et al., 2023) and MineDojo (Fan et al., 2022) extended LLM agents to 3D Minecraft, but the heavy per-game engineering they require limits cross-game generality. The current VLM-agent generation (Li et al., 2025; Bai et al., 2026; Wang et al., 2025b; Tan et al., 2025; Magne et al., 2026) controls GUI or keyboard-mouse interfaces directly across diverse 3D worlds: Game-TARS (Wang et al., 2025b) is pre-trained on over 500B tokens of multimodal gameplay; NitroGen (Magne et al., 2026) on hours of gameplay video across titles; and Lumine (Tan et al., 2025) executes hours-long real-time missions in 3D environments. These agents motivate, but also outpace, the standardized cross-game evaluation infrastructure we provide.

Reflection and Self-improvement. Most game-agent benchmarks report a single-shot score, obscuring whether and how fast an agent improves with repeated interaction. Reflection-based methods address this without weight updates by accumulating natural-language summaries of past experience. Reflexion (Shinn et al., 2023) converts episode feedback into verbal self-critiques; Self-Refine (Madaan et al., 2023) iteratively rewrites model outputs; ExpeL (Zhao et al., 2024) extracts task-level insights; Voyager’s skill library (Wang et al., 2023) accumulates reusable code skills inside Minecraft; and GameVerse (Zhang et al., 2026) reports a single with-vs-without reflection comparison per game. Our work aligns with the recently articulated heuristic learning paradigm (Weng, 2026), which views LLM-driven self-improvement as a learning process operating on explicit artifacts (prompts, code, memory) rather than weights. Our IDC harness instantiates this paradigm in three concrete ways: (i) reflection runs for multiple rounds, producing a full score trajectory rather than a before-after comparison; (ii) the reflector is itself an LLM with tool-use that decides what to inspect rather than executing a fixed-template script; and (iii) we test the resulting skill on held-out task variants per game, revealing differences in skill style that single-number metrics hide.

Game Description Evaluation
\rowcolorgray!15     Solo
ObstacleRun3D A 3D parkour game where the agent navigates to a finish line while avoiding physical obstacles. , where : agent pos., : start, : target
ObstacleRun2D A 2D side-scrolling platformer where the agent must reach the end of a linear level. , where : agent pos., : start, : target
LastStand A platform survival game where the agent must avoid hazards and falling off. , where : time survived, : max duration
MonsterShoot A survival shooting game to locate and eliminate hostile entities while avoiding damage. , where : effective damage, : total enemy health
SceneEscape A scene-based puzzle game requiring the completion of NPC-assigned tasks to escape. , where : completed tasks, : total tasks
CueChase A third-person exploration game to locate and activate hidden triggers across the map. , where : activated triggers, : total triggers
SoloCraft A logistics game where the agent collects, prepares, and delivers items to fulfill orders. , where : fulfilled value, : target value
\rowcolorgray!15     PvP
SkyDuel A direct 1v1 combat game where the agent must engage and defeat an opponent. , where : agent remaining health, : agent max health
CrystalGuard A symmetric attack-and-defend game to destroy the opponent’s crystal core while protecting one’s own. , where : own core health, : core max health
MidlineClash A competitive logistics game where two agents race to fulfill resource orders in a shared environment. , where : agent score, : target score
\rowcolorgray!15     Cooperation
SharedFloor A symmetric cooperative game where agents share a space and capabilities to fulfill delivery orders. , where : fulfilled value, : team target value
HandoffRun An asymmetric cooperative game requiring agents to pass items across restricted areas based on distinct roles. , where : fulfilled value, : team target value
Table 1: Summary of 12 Interactive Games.

3 OmniGameArena

We introduce OmniGameArena, a suite of twelve custom Unreal Engine (UE5) games spanning Solo, PvP, and Coop regimes (Section˜3.1) to systematically evaluate distinct capability axes of vision-based game agents using robust, continuous progress metrics. Furthermore, to address the critical challenge of evaluation contamination, we detail proactive data avoidance strategies and rigorous empirical analysis (Section˜3.2) to ensure the integrity and novelty of our benchmark.

3.1 Game Suite and Evaluation Metrics

The OmniGameArena suite comprises twelve visually rich and physically complex environments developed in Unreal Engine 5. Each game is purposefully designed to isolate and test specific subsets of embodied capabilities, ranging from solitary spatial reasoning to complex multi-agent cooperation, as shown in Figure˜2. Game progress is uniformly normalized to a continuous scale of , providing a consistent metric across highly diverse tasks. We provide a brief description and its evaluation protocol for each game, as demonstrated in Section˜2.

3.2 Contamination Avoidance

We adopt a proactive approach to mitigate data contamination during the benchmark design phase. We begin by conducting a web-exposure audit to search for exact game names, task phrases, rule descriptions, and scoring events prior to release, ensuring these specific elements are strictly excluded from the benchmark. To guarantee environment novelty, we construct entirely new games using Unreal Engine 5 (UE5). For visual content, we utilize a mixture of bespoke and off-the-shelf UE5 marketplace assets. However, the combination of assets, the design of the level geometry, the execution order of scripts, and the criteria for success are uniquely designed, ensuring the evaluation scenarios cannot be memorized from pre-training data.

媒体内容 · 前往原文查看
Benchmark Games(#) Recog.(%) Mech. (%)
BALROG 06 066.7 100.0
LMGame-Bench 06 100.0 100.0
ORAK 12 100.0 100.0
OmniGameArena 12 000.0 050.0
Table 2: Contamination analysis when given screenshots. Recog.: Percentage of games recognized; Mech.: Percentage of games where underlying mechanics were successfully described.

Contamination Analysis. To empirically verify the effectiveness of our avoidance strategies, we conduct a contamination analysis focusing on visual novelty and rule leakage. First, we provide a representative model (e.g., Gemini) with screenshots of games to assess whether it can recognize the game name, confirming the visual novelty of the tasks. Second, we evaluate whether the model can successfully describe the underlying mechanics of the games based purely on visual inputs, which tests for the leakage of memorized game rules. As shown in Table 2, we compare existing benchmarks (Paglieri et al., 2024; Park et al., 2025; Hu et al., 2025) against our proposed OmniGameArena across these two tests. The results clearly demonstrate that the games within the existing benchmarks are highly recognizable, allowing the model to retrieve their mechanics directly from memory. In contrast, OmniGameArena exhibits a recognition rate and significantly reduced mechanics leakage (). These results suggest that OmniGameArena substantially mitigates the risk of pre-training contamination, reducing the influence of memorized priors on agent evaluation.

Refer to caption
Figure 3: Overview of the Improvement Dynamics Curve (IDC) harness. The experience acquisition module (left) runs episodes under the current skill. The reflection module (center) reads both the new trajectories and the persistent state (notebook and prior skill), then runs four autonomous stages (Explore, Diagnose, Validate, Distill) to produce a refined skill. The persistent module (right) stores the experience notebook, validated skills, and per-round score curve, which seed the next round.

4 Game Agent Harness

OmniGameArena specifies the games; the harness specifies how an agent plays them. The harness has two layers: a per-episode loop (§4.1) that drives any agent during cold-start runs, and a reflective outer loop (§4.2) whose round-level scores form the Improvement Dynamics Curve studied in §5.3.

4.1 Per-Episode Loop

OmniGameArena exposes a Gym-like interface. At step the harness receives observation , where is the RGB frame at resolution and is its capture timestamp. The agent emits an action (a chunked keyboard-mouse action for VLMs), which the engine executes before returning ; the loop terminates on done.

Bounded visual-action history.

For VLM agents, the harness maintains a sliding window of the last observation-response pairs:

where is the raw VLM response at step . At step , the VLM is prompted with system instructions, an optional skill prompt , the history , and the current frame , for up to images per call. The current frame is appended to history only after is returned.

4.2 Improvement Dynamics Curve

IDC wraps the per-episode loop with a reflective outer loop organized into three modules (Figure 3): an experience acquisition module that runs episodes under the current skill, a reflection module that converts the resulting trajectories into a refined skill, and a persistent module that carries state across rounds. We deliberately keep the loop lightweight yet functionally complete: a small fixed tool surface gives the reflector full agentic autonomy without prescribing an inspection script.

Let denote a VLM with frozen weights . At round the agent is conditioned on a skill prompt and acts according to ; the skill is fixed throughout the round.

Experience acquisition. Each round contains episodes, yielding trajectories and round score

Round uses an empty skill () and serves as the cold-start baseline.

Agent ObstacleRun2D ObstacleRun3D LastStand MonsterShoot SceneEscape CueChase SoloCraft
[Uncaptioned image]   Claude Opus 4.7 (Anthropic, 2026b) \cellcolorsecond \cellcolorthird \cellcolorthird
[Uncaptioned image]   Claude Opus 4.6 (Anthropic, 2026a) \cellcolorsecond \cellcolorsecond \cellcolorbest \cellcolorsecond
[Uncaptioned image]   Claude Sonnet 4.6 (Anthropic, 2026c) \cellcolorbest
[Uncaptioned image]   GPT-5.5 (OpenAI, 2026b) \cellcolorbest \cellcolorbest \cellcolorsecond \cellcolorbest \cellcolorthird \cellcolorbest
[Uncaptioned image]   GPT-5.4 (OpenAI, 2026a) \cellcolorsecond
[Uncaptioned image]   Gemini 3.1 Pro Preview (Google, 2026b) \cellcolorthird \cellcolorbest \cellcolorthird \cellcolorsecond
[Uncaptioned image]   Gemini 3.1 Flash-Lite Preview (Google, 2026b) \cellcolorthird
[Uncaptioned image]   Kimi K2.5 (Moonshot AI, 2026) \cellcolorthird
[Uncaptioned image]   Qwen3.5-397B-A17B (Qwen Team, 2026)
[Uncaptioned image]   Qwen3.5-122B-A10B (Qwen Team, 2026)
[Uncaptioned image]   NitroGen (gamepad) (Magne et al., 2026)
[Uncaptioned image]   Open-P2P (kbd-mouse) (Yue et al., 2026)
Table 3: Solo cold-start scores. Top 3 per column: red orange yellow.

Reflection. After round , the reflector refines the skill autonomously, reading the new trajectories together with the persistent state (notebook and prior skill) and deciding for itself which content to inspect, how many tool calls to spend, and when to terminate. The refinement proceeds in four stages. The Explore stage exposes the round’s per-episode trajectories through sandboxed read-only tools (list_dir, read_text, read_image, grep); the reflector chooses what to inspect rather than executing a fixed script. The Diagnose stage commits an explicit list of failure modes via submit_diagnosis, separating causes from prescriptions. The Validate stage proposes an updated skill and calls validate_skill, an independent LLM judge that rejects proposals which memorize map content, or contradict the diagnosis; the reflector iterates up to five times before committing. The Distill stage finalizes the accepted skill and, optionally, edits the notebook with durable observations. The combination of agentic autonomy within a small, fixed tool surface (read, diagnose, judge, write) is what keeps the framework lightweight without sacrificing functional completeness.

Persistent module. Three artifacts carry state across rounds. The experience notebook is a factual log written by the reflector for its own future use (e.g., “round episode step : event”), capped at tokens, edited rather than appended, and never shown to the player. The validated skills comprise the skill prompt used in the current round (player-visible, capped at tokens of cue-to-response heuristics) together with the best-skill cache used for rollback. The curve is the score sequence accumulated so far.

Best-skill rollback. Reflection is not monotone. When a round’s score drops sharply below the best seen so far (, ), the harness resets the next round’s starting prompt to and resumes reflection from there, guarding against catastrophic skill drift.

The curve. Iterating for rounds yields , the Improvement Dynamics Curve of the (agent, game) pair. Two models with the same final can produce very different curves (early vs. late convergence, monotone vs. oscillating); the curve, rather than any single round, is the measurement object in §5.3.

5 Experiments

We report results in three blocks. §5.1 fixes the evaluation protocol, the set of agents, and the scoring conventions used throughout. §5.2 reports the main cold-start leaderboard, in which every agent plays every game once with no prior trajectories and no provided skill, spanning the seven Solo, three PvP, and two Coop games. §5.3 relaxes the cold-start constraint via the Improvement Dynamics Curve (IDC), an agentic self-reflection protocol in which the same model alternates between playing the game and rewriting its own skill (a natural-language summary of game-specific strategy) for rounds. We use IDC to characterize (i) how each model improves on the original task, and (ii) whether the learned skill transfers to three held-out task variants per game.

5.1 Experimental Settings

Agents. We evaluate twelve agents grouped into three classes so their strengths and limitations are not averaged away. (a) Commercial VLMs (API-only): three Claude models (Opus 4.7 (Anthropic, 2026b), Opus 4.6 (Anthropic, 2026a), Sonnet 4.6 (Anthropic, 2026c)), two OpenAI models (GPT-5.5 (OpenAI, 2026b) and GPT-5.4 (OpenAI, 2026a)), two Gemini models (3.1 Pro Preview and 3.1 Flash-Lite Preview (Google, 2026b)), and Kimi K2.5 (Moonshot AI, 2026). (b) Open-weight VLMs: two Qwen3.5 mixture-of-experts checkpoints, 397B-A17B and 122B-A10B (Qwen Team, 2026), served locally behind an OpenAI-compatible endpoint. (c) Specialized game policies: NitroGen (Magne et al., 2026), which consumes a single frame and emits a chunk of 18 gamepad actions (21-dim each), and Open-P2P (Yue et al., 2026), which consumes a 200-frame keyboard-mouse history and emits one action per frame. Both run with the native real-time protocols from their original papers. All agents are evaluated on the same OmniGameArena real-time environment with matched configuration. VLMs use OmniGameArena’s chunked keyboard-mouse adapter, while the specialized policies are routed through their native interfaces unchanged, so each system is evaluated at the operating point it was designed for.

Evaluation protocol. Each (agent, game) cell is evaluated over episodes; PvP cells additionally cover every pairwise matchup in the agent pool. OmniGameArena natively runs in real time, but commercial VLM API calls suffer from network jitter that is orthogonal to model capability. We therefore evaluate under two clock modes that both pause the environment during inference: Paused Decision Quality (PDQ) freezes the environment for the full inference call and treats decision time as free, isolating pure decision quality; Latency-Controlled Real-Time (LCRT) additionally idles for the server-reported inference time before each action is applied, charging agents for on-device latency while excluding network round-trip noise. Main-table results use PDQ and LCRT is reported in Appendix A.

Refer to caption
Figure 4: PvP win rates of Player 1 (row) against Player 2 (column) per game over all pairings.

5.2 Cold-start Leaderboard

Solo games. Table 3 reports the seven Solo games. Three patterns stand out. (1) No single model dominates. Leadership rotates across games: GPT-5.5 leads four of seven games (ObstacleRun2D, LastStand, SceneEscape, SoloCraft), Claude Opus 4.6 wins CueChase by a wide margin ( vs. next ), and Gemini 3.1 Pro leads MonsterShoot ( vs. next ). (2) Newer is not always better. Claude Opus 4.6 outperforms Opus 4.7 on five of seven Solo games, and GPT-5.4 exceeds GPT-5.5 on SceneEscape, indicating that capability ranking is task-specific rather than monotone in release order. (3) Open-weight VLMs and specialized policies fail to transfer. Both Qwen3.5 MoE checkpoints score below on every game and exactly on several, while NitroGen (gamepad) and Open-P2P (keyboard-mouse) collapse to near zero on all but a handful of games. This confirms that OmniGameArena’s task diversity lies well outside the training distribution of policies optimized for narrow single-game.

PvP games. Figure 4 reports Player 1 win rates over all pairings (diagonal omitted). SkyDuel and CrystalGuard show a clean dominance hierarchy that tracks the Solo leaderboard: GPT-5.5 and Gemini 3.1 Pro win against nearly every opponent, while both Qwen3.5 variants lose nearly every matchup. MidlineClash is non-transitive: Kimi K2.5 beats Claude Opus 4.6 in all five Player 1 matches () even though Claude is the stronger Solo agent, and Claude wins decisively only against Qwen3.5. This indicates that MidlineClash rewards game-specific tactics that do not align with the Solo capability ranking.

Coop games. Table 4 reports team scores when two copies of the same model cooperate. The leaderboard mirrors Solo: GPT-5.5 leads both games, Gemini 3.1 Pro is a close second, and Claude Opus 4.6 follows. Two findings extend beyond Solo. First, the gap between commercial and open-weight VLMs widens: both Qwen3.5 checkpoints score exactly on both games, failing to complete a single shared order or hand-off. Second, even the strongest model reaches only on SharedFloor and on HandoffRun, leaving substantial headroom and indicating that LLM-LLM coordination is an unsolved capability gap rather than a saturated benchmark dimension.

Agent SharedFloor HandoffRun
[Uncaptioned image]  Claude Opus 4.7 (Anthropic, 2026b)
[Uncaptioned image]  Claude Opus 4.6 (Anthropic, 2026a) \cellcolorthird \cellcolorthird
[Uncaptioned image]  Claude Sonnet 4.6 (Anthropic, 2026c)
[Uncaptioned image]  GPT-5.5 (OpenAI, 2026b) \cellcolorbest \cellcolorbest
[Uncaptioned image]  GPT-5.4 (OpenAI, 2026a)
[Uncaptioned image]  Gemini 3.1 Pro Preview (Google, 2026b) \cellcolorsecond \cellcolorsecond
[Uncaptioned image]  Gemini 3.1 Flash-Lite Preview (Google, 2026a)
[Uncaptioned image]  Kimi K2.5 (Moonshot AI, 2026)
[Uncaptioned image]  Qwen3.5-397B-A17B (Qwen Team, 2026)
[Uncaptioned image]  Qwen3.5-122B-A10B (Qwen Team, 2026)
Table 4: Coop team scores on two cooperative games, where two copies of the same model must coordinate to complete a shared task.
LastStand SharedFloor
Agent origin var1 var2 var3 origin var1 var2 var3
[Uncaptioned image]  Claude Opus 4.6 \cellcolorgainstrong \cellcolorgainstrong \cellcolorlossstrong \cellcolorgainmild \cellcolorgainmid \cellcolorgainstrong \cellcolorgainmid \cellcolorgainmid
[Uncaptioned image]  Claude Opus 4.7 \cellcolorgainstrong \cellcolorlossmid \cellcolorlossstrong \cellcolorlossmid \cellcolorgainstrong \cellcolorgainmid \cellcolorgainmid \cellcolorgainmid
[Uncaptioned image]  GPT-5.5 \cellcolorgainstrong \cellcolorgainstrong \cellcolorgainstrong \cellcolorgainmild \cellcolorgainmild \cellcolorgainmid \cellcolorgainmild \cellcolorgainmild
[Uncaptioned image]  Gemini 3.1 Pro Preview \cellcolorgainstrong \cellcolorgainstrong \cellcolorlossmid \cellcolorgainmid \cellcolorgainmid \cellcolorgainmid \cellcolorgainmild \cellcolorgainmild
Table 5: Transfer of IDC best skill across the origin split and three unseen variants. Each cell shows the gain followed by the -change relative to . is the no-skill baseline on that split; applies the skill that gave the highest mean score during the 10-round run.

5.3 Improvement Dynamics Curve

Setup. We run IDC on LastStand (Solo) and SharedFloor (Coop) for complementary skill coverage: LastStand stresses reactive control as tiles progressively fall away, while SharedFloor combines explicit task rules with multi-agent coordination. We exclude PvP to keep IDC gains attributable to the reflector rather than to opponent behavior. Four top-performing agents from the cold-start leaderboard participate: Claude Opus 4.6, Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro. Each agent completes rounds of episodes under PDQ; Round  reuses the cold-start baseline from §5.2. The best skill from each (model, game) run is then reapplied to three held-out task variants per game.

Refer to caption
Figure 5: IDC curves: per-round mean episode score across 10 reflection rounds for four agents on LastStand (left) and SharedFloor (right). Horizontal lines mark each agent’s R0 (no-skill PDQ baseline) for reference; deviations above the line indicate skill-driven improvement. Each round aggregates 5 episodes per agent.

5.3.1 Improvement on the original task

Figure 5 reports per-round mean score. Through multi-round agentic reflection, every model on both games discovers skills that improve over its Round 0 baseline. On LastStand, best-round gains range from to in survival fraction ( to relative). On SharedFloor, best-round teams complete to additional orders per episode, a substantial improvement in coordination throughput. However, peak performance is typically reached mid-curve rather than at Round 10: on LastStand the two Opus models lose to between best and final round, while GPT-5.5 and Gemini retain almost all gains. This best-vs-final gap justifies the best-skill rollback mechanism described in §3.

5.3.2 Transfer to held-out variants

Table 5 reports the gain of each best skill on the origin split and three held-out variants per game. SharedFloor transfers universally ( positive across variants). Gains range from to , corresponding to roughly to additional completed orders per episode under the best skill. The skills encode coordination heuristics rather than spatial memory: agents are guided to observe the teammate’s position and currently held item before committing to an order (to avoid both players picking up the same item), to discard duplicates at the trash bin when this happens, to keep distance from the teammate’s workstation to prevent crowding, and to maintain a visible division of labor across stations. Because the variants only change workbench and item placements while preserving coordination rules, these behavioral heuristics transfer without modification. LastStand transfer depends on skill style, not origin gain magnitude. The origin drops tiles one at a time, so three of four models converge to a “find a safe tile and stay” policy that exploits the single-tile structure. var1 (different seed, same mechanics) leaves the structure intact and transfers positively for three of four models. var2 (cluster drops of multiple connected tiles) removes the safe pocket the skills relied on: both Opus models collapse ( and ), while GPT-5.5 gains . Skill inspection (Appendix C) explains the divergence: the Opus and Gemini skills converge on movement-minimizing policies (e.g., “stand still unless your tile is red”, “never chain forward steps”), which are optimal under single-tile drops but maladaptive when a cluster wipes out the safe pocket. GPT-5.5’s skill instead encodes a “move briefly, then reassess” loop that adapts to either mechanic. var3 (tiles that track the player) eliminates static safety entirely; transfer drops to small or negative gains for all four models. GPT-5.5 is the only model with positive transfer on all three variants yet has the smallest origin gain (); Opus 4.7 gains on origin but transfers negatively on every variant. This dissociation between origin gain and transferability is the central finding of the variant experiment.

6 Conclusion

We introduced OmniGameArena, a benchmark of twelve newly built UE5 real-time games spanning Solo, PvP, and Coop, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness that produces multi-round self-improvement trajectories. Beyond single-round leaderboard scores, the IDC exposes two additional observables for each (agent, game) pair: how the score evolves across reflection rounds, and how the learned skill behaves on held-out task variants.

Limitations

IDC scope. Due to compute constraints, our IDC experiments cover only two environments (LastStand and SharedFloor), each with three held-out variants, and four agents from the cold-start leaderboard. Scaling IDC to additional games, variants, and models is an extension.

Single-skill format. Our reflector maintains a single bounded skill prompt that is replaced each round, rather than a growing library of skills in the style of Voyager. Library-based extensions are orthogonal to the round-by-round refinement.

Shared model for player and reflector. Each agent uses the same underlying model as both player and reflector. Whether asymmetric setups (e.g., a smaller player paired with a stronger reflector) yield different improvement is untested.

References

  • Anthropic (2026a) Anthropic. 2026a. Introducing Claude Opus 4.6. https://www.anthropic.com/news/claude-opus-4-6.
  • Anthropic (2026b) Anthropic. 2026b. Introducing Claude Opus 4.7. https://www.anthropic.com/news/claude-opus-4-7.
  • Anthropic (2026c) Anthropic. 2026c. Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/claude-sonnet-4-6.
  • Bai et al. (2026) Hao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar, and Spencer Whitehead. 2026. Webgym: Scaling training environments for visual web agents with realistic tasks. arXiv preprint arXiv:2601.02439.
  • Fan et al. (2022) Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. 2022. Minedojo: Building open-ended embodied agents with internet-scale knowledge. volume 35, pages 18343–18362.
  • Feng et al. (2023) Xidong Feng, Yicheng Luo, Ziyan Wang, Hongrui Tang, Mengyue Yang, Kun Shao, David Mguni, Yali Du, and Jun Wang. 2023. Chessgpt: Bridging policy learning and language modeling. Advances in Neural Information Processing Systems, 36:7216–7262.
  • Google (2026a) Google. 2026a. Introducing Gemini 3.1 Flash-Lite. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-lite/.
  • Google (2026b) Google. 2026b. Introducing Gemini 3.1 Pro. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/.
  • Hausknecht et al. (2020) Matthew Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, and Xingdi Yuan. 2020. Interactive fiction games: A colossal adventure. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7903–7910.
  • Hu et al. (2025) Lanxiang Hu, Mingjia Huo, Yuxuan Zhang, Haoyang Yu, Eric P Xing, Ion Stoica, Tajana Rosing, Haojian Jin, and Hao Zhang. 2025. lmgame-bench: How good are llms at playing games? arXiv preprint arXiv:2505.15146.
  • Hu et al. (2024) Lanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang, Ion Stoica, Haojian Jin, and Hao Zhang. 2024. Gamearena: Evaluating llm reasoning through live computer games. arXiv preprint arXiv:2412.06394.
  • Huang et al. (2024) Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael R Lyu. 2024. How far are we on the decision-making of llms? evaluating llms’ gaming ability in multi-agent environments. arXiv preprint arXiv:2403.11807.
  • Küttler et al. (2020) Heinrich Küttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel. 2020. The nethack learning environment. Advances in Neural Information Processing Systems, 33:7671–7684.
  • Li et al. (2025) Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, and Yitao Liang. 2025. Jarvis-vla: Post-training large-scale vision language models to play visual games with keyboards and mouse. In Findings of the Association for Computational Linguistics: ACL 2025, pages 17878–17899.
  • Lin et al. (2025) Mingxian Lin, Wei Huang, Yitang Li, Chengjie Jiang, Kui Wu, Fangwei Zhong, Shengju Qian, Xin Wang, and Xiaojuan Qi. 2025. Embrace-3k: Embodied reasoning and action in complex environments. arXiv preprint arXiv:2507.10548.
  • Luo et al. (2025) Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu, Liyao Tang, Yuqi Liu, Ying-Cong Chen, Shengju Qian, Xin Wang, and Yang You. 2025. V-reasonbench: Toward unified reasoning benchmark suite for video generation models. arXiv preprint arXiv:2511.16668.
  • Madaan et al. (2023) Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2023. Self-refine: Iterative refinement with self-feedback. Advances in neural information processing systems, 36:46534–46594.
  • Magne et al. (2026) Loïc Magne, Anas Awadalla, Guanzhi Wang, Yinzhen Xu, Joshua Belofsky, Fengyuan Hu, Joohwan Kim, Ludwig Schmidt, Georgia Gkioxari, Jan Kautz, and 1 others. 2026. Nitrogen: An open foundation model for generalist gaming agents. arXiv preprint arXiv:2601.02427.
  • Moonshot AI (2026) Moonshot AI. 2026. Kimi k2.5: Visual agentic intelligence. https://www.kimi.com/blog/kimi-k2-5.
  • OpenAI (2026a) OpenAI. 2026a. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4/.
  • OpenAI (2026b) OpenAI. 2026b. Introducing GPT-5.5. https://openai.com/index/introducing-gpt-5-5/.
  • Paglieri et al. (2024) Davide Paglieri, Bartłomiej Cupiał, Samuel Coward, Ulyana Piterbarg, Maciej Wolczyk, Akbir Khan, Eduardo Pignatelli, Łukasz Kuciński, Lerrel Pinto, Rob Fergus, and 1 others. 2024. Balrog: Benchmarking agentic llm and vlm reasoning on games. arXiv preprint arXiv:2411.13543.
  • Park et al. (2025) Dongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim, Keon Lee, Jonghyun Lee, Inkyu Park, Byeong-Uk Lee, Jaeyoung Hwang, Jaewoo Ahn, and 1 others. 2025. Orak: A foundational benchmark for training and evaluating llm agents on diverse video games. arXiv preprint arXiv:2506.03610.
  • Qwen Team (2026) Qwen Team. 2026. Qwen3.5: A Native Multimodal Foundation Model for Efficiency. https://qwen.ai/blog?id=qwen3.5.
  • Shinn et al. (2023) Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in neural information processing systems, 36:8634–8652.
  • Tan et al. (2024) Weihao Tan, Ziluo Ding, Wentao Zhang, Boyu Li, Bohan Zhou, Junpeng Yue, Haochong Xia, Jiechuan Jiang, Longtao Zheng, Xinrun Xu, and 1 others. 2024. Towards general computer control: A multimodal agent for red dead redemption ii as a case study. arXiv preprint arXiv:2403.03186, 1(2).
  • Tan et al. (2025) Weihao Tan, Xiangyang Li, Yunhao Fang, Heyuan Yao, Shi Yan, Hao Luo, Tenglong Ao, Huihui Li, Hongbin Ren, Bairen Yi, and 1 others. 2025. Lumine: An open recipe for building generalist agents in 3d open worlds. arXiv preprint arXiv:2511.08892.
  • Tsai et al. (2023) Chen Feng Tsai, Xiaochen Zhou, Sierra S Liu, Jing Li, Mo Yu, and Hongyuan Mei. 2023. Can large language models play text games well? current state-of-the-art and open questions. arXiv preprint arXiv:2304.02868.
  • Wang et al. (2023) Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models, 2023. URL https://arxiv. org/abs/2305.16291, 2(11).
  • Wang et al. (2025a) Xinyu Wang, Bohan Zhuang, and Qi Wu. 2025a. Are large vision language models good game players? arXiv preprint arXiv:2503.02358.
  • Wang et al. (2025b) Zihao Wang, Xujing Li, Yining Ye, Junjie Fang, Haoming Wang, Longxiang Liu, Shihao Liang, Junting Lu, Zhiyong Wu, Jiazhan Feng, and 1 others. 2025b. Game-tars: Pretrained foundation models for scalable generalist multimodal game agents. arXiv preprint arXiv:2510.23691.
  • Weng (2026) Jiayi Weng. 2026. Learning beyond gradients. https://trinkle23897.github.io/learning-beyond-gradients/. Blog post.
  • Wu et al. (2023) Yue Wu, Xuan Tang, Tom M Mitchell, and Yuanzhi Li. 2023. Smartplay: A benchmark for llms as intelligent agents. arXiv preprint arXiv:2310.01557.
  • Yue et al. (2026) Yuguang Yue, Irakli Salia, Samuel Hunt, Chris Green, Wenzhe Shi, and Jonathan J Hunt. 2026. Scaling behavior cloning improves causal reasoning: An open model for real-time video game playing. arXiv preprint arXiv:2601.04575.
  • Zhang et al. (2025) Alex L Zhang, Thomas L Griffiths, Karthik R Narasimhan, and Ofir Press. 2025. Videogamebench: Can vision-language models complete popular video games? arXiv preprint arXiv:2505.18134.
  • Zhang et al. (2026) Kuan Zhang, Dongchen Liu, Qiyue Zhao, Jinkun Hou, Xinran Zhang, Qinlei Xie, Miao Liu, and Yiming Li. 2026. Gameverse: Can vision-language models learn from video-based reflection? arXiv preprint arXiv:2603.06656.
  • Zhao et al. (2024) Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19632–19642.
  • Zheng et al. (2025) Xiangxi Zheng, Linjie Li, Zhengyuan Yang, Ping Yu, Alex Jinpeng Wang, Rui Yan, Yuan Yao, and Lijuan Wang. 2025. V-mage: A game evaluation framework for assessing vision-centric capabilities in multimodal large language models. arXiv preprint arXiv:2504.06148.
  • Zhu et al. (2026) Lingting Zhu, Shengju Qian, Haidi Fan, Jiayu Dong, Zhenchao Jin, Siwei Zhou, Gen Dong, Xin Wang, and Lequan Yu. 2026. Assetformer: Modular 3d assets generation with autoregressive transformer. arXiv preprint arXiv:2602.12100.

Appendix A Latency-Controlled Real-Time Eval

We further evaluate a subset of agents under a Latency-Controlled Real-Time (LCRT) protocol. In PDQ, the simulator is paused while the model is thinking, and the returned action is executed immediately from the observed state. In LCRT, the simulator is likewise paused during the wall-clock model call, but the model’s measured decision latency is then injected back into the game timeline before the action is executed. Thus an action predicted from observation is applied after approximately seconds of simulated game time, where is the model’s decision latency. LCRT requires a reliable estimate of pure model inference time. We therefore report LCRT only for the four models whose backends expose usable model-side timing signals in our implementation: Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.5, and GPT-5.4. Other agents are excluded because their available timings fold in client-side overheads such as request queuing, network transfer, retries, or wrapper latency; injecting such end-to-end wall-clock measurements would confound model latency with infrastructure latency and make the comparison unfair.

Solo Games. We focus the LCRT analysis on tasks where latency is expected to affect either the evolving game state or the available time budget, and accordingly report LastStand and MonsterShoot as dynamic real-time tasks and SoloCraft as a time-budgeted interaction task. As shown in Table˜6, the three tasks reveal distinct sensitivity patterns. MonsterShoot is the most latency-sensitive: all four models drop, most steeply GPT-5.5 (), consistent with its need for continuous aiming and target tracking. LastStand behaves differently: its optimal policy is nearly stationary, waiting on a safe tile and moving only when one’s own tile is about to fall, so latency is not necessarily harmful here, and taking fewer, later actions can even avoid fatal missteps. This may explain why Claude Opus 4.6 and GPT-5.4 even improve under LCRT and Claude Sonnet 4.6 stays essentially flat, while GPT-5.5 is the only model that degrades. SoloCraft is not a reactive-control task in the same sense; here latency mainly consumes the episode time budget and reduces interaction throughput, an effect we quantify below.

Agent LastStand MonsterShoot SoloCraft
[Uncaptioned image]   Claude Opus 4.6 (Anthropic, 2026a) +0.101 -0.182 -0.148
[Uncaptioned image]   Claude Sonnet 4.6 (Anthropic, 2026c) +0.010 -0.050 -0.052
[Uncaptioned image]   GPT-5.5 (OpenAI, 2026b) -0.092 -0.408 -0.200
[Uncaptioned image]   GPT-5.4 (OpenAI, 2026a) +0.105 -0.188 -0.044
Table 6: LCRT results on selected Solo games. Each cell reports mean with standard deviation as subscript over episodes, with denoting the change relative to the corresponding PDQ result.
Refer to caption
Figure 6: PvP win rates of Player 1 (row) against Player 2 (column) on MidlineClash under latency control setting.

PvP Games. Figure˜6 reports Player 1 win rates on MidlineClash under LCRT for the four commercial VLMs. They differ from the PDQ heatmap (Figure˜4), but the difference reflects the change of clock mode rather than a change in model ability. Charging decision latency against the game clock leaves far less game time per move, so each player completes only about actions per game under LCRT against the -action PDQ budget, and matches compress into low-scoring, frequently drawn games. The effect is clearest in the GPT-5.5-vs-Opus 4.6 pairing, the one matchup both protocols share: GPT-5.5 swept it – under PDQ by wide margins (e.g. –, –, –), whereas under LCRT the same pairing yields GPT-5.5 wins, Opus wins, and draws, with narrow scorelines such as –, –, and – and the pairing’s average per-player score falling from to . GPT-5.5 still wins the pairing and still posts the highest average Player-1 win rate (, versus for the weakest, GPT-5.4), even though it carries by far the largest inference latency ( s of pure model time against – s for the others). Because both sides pay the same delay, latency in symmetric play mainly compresses margins and amplifies single-game randomness rather than re-ranking the agents, a much milder effect than on the throughput tasks below.

Agent SharedFloor
[Uncaptioned image]   Claude Opus 4.6 (Anthropic, 2026a) -0.104
[Uncaptioned image]   Claude Sonnet 4.6 (Anthropic, 2026c) -0.120
[Uncaptioned image]   GPT-5.5 (OpenAI, 2026b) -0.320
[Uncaptioned image]   GPT-5.4 (OpenAI, 2026a) -0.044
Table 7: LCRT results on the cooperative game SharedFloor. Each cell reports mean with standard deviation as subscript over episodes, with denoting the change relative to the corresponding PDQ result.

Cooperative Game. We further report the cooperative task SharedFloor, in which two instances of the same model coordinate to complete shared orders before a fixed match deadline (Table˜7). Every model degrades under LCRT, and the loss scales with the action budget: because LCRT charges each model’s decision latency against the match clock, the slowest agent completes the fewest interactions. GPT-5.5, the slowest, takes the largest absolute drop, from under PDQ to (); but having started far ahead, it still ties Opus 4.6 for the best LCRT score rather than collapsing. GPT-5.4, the fastest, retains the most actions and changes the least (). We quantify this action-budget account next.

Action budget versus per-action efficiency. To see why scores fall on the throughput tasks, we measure for each agent both its actions per episode and its score per action under the two protocols on SoloCraft (Table˜8) and SharedFloor (Table˜9). The action budget shrinks monotonically with model speed: GPT-5.5, the slowest, completes the fewest actions, only about of its PDQ actions per episode on SoloCraft and versus on SharedFloor, while the faster agents retain far more. The drop is overwhelmingly a budget effect: score per action is largely preserved between PDQ and LCRT, most clearly for GPT-5.5 (essentially unchanged on SoloCraft, only mildly lower on SharedFloor), so each agent’s score falls roughly in proportion to its smaller action count, e.g. GPT-5.5 on SoloCraft drops for actions. This account is specific to interaction-throughput tasks, where score accumulates with the number of useful interactions before a fixed deadline; it does not extend to the dynamic tasks, where latency acts through reaction timing rather than throughput and is not even monotone, helping in LastStand, whose near-stationary optimal policy rewards fewer and later actions, but hurting in MonsterShoot through missed and mistimed shots. Together these regimes show that latency is a first-class evaluation axis whose effect is task-dependent, taxing single-agent throughput, compressing rather than re-ranking symmetric play, and even helping near-stationary survival, so PDQ margins do not transfer directly to real-time deployment.

Actions / episode Score / action
Agent PDQ LCRT PDQ LCRT
[Uncaptioned image]   Claude Opus 4.6 (Anthropic, 2026a) 43 17 0.0053 0.0047
[Uncaptioned image]   Claude Sonnet 4.6 (Anthropic, 2026c) 43 19 0.0029 0.0038
[Uncaptioned image]   GPT-5.5 (OpenAI, 2026b) 43 8 0.0059 0.0065
[Uncaptioned image]   GPT-5.4 (OpenAI, 2026a) 43 28 0.0020 0.0014
Table 8: SoloCraft action budget. Actions per episode and normalized score per action under PDQ vs. LCRT. Score per action is nearly unchanged, so each agent’s LCRT drop (Table˜6) is almost entirely its smaller action budget.
Actions / episode Score / action
Agent PDQ LCRT PDQ LCRT
[Uncaptioned image]   Claude Opus 4.6 (Anthropic, 2026a) 84 33 0.0018 0.0015
[Uncaptioned image]   Claude Sonnet 4.6 (Anthropic, 2026c) 84 36 0.0018 0.0008
[Uncaptioned image]   GPT-5.5 (OpenAI, 2026b) 84 14 0.0044 0.0034
[Uncaptioned image]   GPT-5.4 (OpenAI, 2026a) 84 55 0.0008 0.0004
Table 9: SharedFloor action budget (actions and team score summed over both players; score normalized). The LCRT drop tracks the shrinking action budget: GPT-5.5, the slowest, completes by far the fewest actions, yet keeps the highest score per action under both protocols, so its loss is a budget effect rather than a per-action collapse.
Refer to caption
Figure 7: Qualitative comparison on Last Stand using GPT-5.5.
Refer to caption
Figure 8: Qualitative comparison on Shared Floor using Gemini-3.1-Pro.

Appendix B Qualitative Comparison with IDC

IDC skill induction improves the agents’ behavior across both environments. Without IDC, the agents tend to make unstable or inefficient decisions: in the survival-oriented task, the agent fails to consistently maintain a safe position, while in the cooperative task, the agents show weaker coordination and complete fewer objectives. After IDC, the learned behaviors become more effective and task-aligned. The agent in the survival task maintains safer positions and achieves a higher score, and the agents in the cooperative task exhibit better coordination and obtain a higher team score. (See Figure 8)

Appendix C Skill Inspection

This appendix lists the best measured skill prompt for each game–model pair in the IDC comparison. It explains the divergence discussed in the main text: LastStand prompts converge on conservative tile-survival behavior, while SharedFloor prompts emphasize cooperative division of labor, station alignment, and order-refresh handling. The scope here is limited to LastStand and SharedFloor; ObstacleRun3D is intentionally excluded.

Appendix D Visualization

Visualization results are shown in Figures 9–20. For each game, we visualize representative trajectories from different models. Each row corresponds to one model or matchup, with five sampled frames illustrating the progression of the episode.

媒体内容 · 前往原文查看
Table 10: Best skill prompt for LastStand using claude-opus-4-6.
媒体内容 · 前往原文查看
Table 11: Best skill prompt for LastStand using claude-opus-4-7.
媒体内容 · 前往原文查看
Table 12: Best skill prompt for LastStand using gemini-3.1-pro-preview.
媒体内容 · 前往原文查看
Table 13: Best skill prompt for LastStand using gpt-5.5.
媒体内容 · 前往原文查看
Table 14: Best skill prompt for SharedFloor using claude-opus-4-6.
媒体内容 · 前往原文查看
Table 15: Best skill prompt for SharedFloor using claude-opus-4-7.
媒体内容 · 前往原文查看
Table 16: Best skill prompt for SharedFloor using gemini-3.1-pro-preview.
媒体内容 · 前往原文查看
Table 17: Best skill prompt for SharedFloor using gpt-5.5.
Refer to caption
Figure 9: Visualization results for cue_chase. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 10: Visualization results for last_stand. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 11: Visualization results for monster_shoot. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 12: Visualization results for obstacle_run_2d. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 13: Visualization results for obstacle_run_3d. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 14: Visualization results for scene_escape. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 15: Visualization results for solo_craft. Each row shows one model, with five sampled frames from the corresponding trajectory.
Refer to caption
Figure 16: Visualization results for handoff_run. Each row shows one cooperative model pair, with five sampled frames from the corresponding episode.
Refer to caption
Figure 17: Visualization results for shared_floor. Each row shows one cooperative model pair, with five sampled frames from the corresponding episode.
Refer to caption
Figure 18: Visualization results for crystal_guard. Each row shows one representative PvP matchup, with five sampled frames from the corresponding match.
Refer to caption
Figure 19: Visualization results for midline_clash. Each row shows one representative PvP matchup, with five sampled frames from the corresponding match.
Refer to caption
Figure 20: Visualization results for sky_duel. Each row shows one representative PvP matchup, with five sampled frames from the corresponding match.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org