Skip to main content
Back to timeline
arXivSource publication:

PlaySuite benchmarks 14 open models on 5,734 open-source games and finds a clear perception-action gap

Related research and updates

Synopsis

The work builds PlaySuite, a benchmark curated from PyWeek and itch.io spanning 5,734 open-source games, together with a unified closed-loop interaction framework, HPC-oriented batched execution, and a Video-LLM-as-a-judge milestone scoring protocol, and evaluates 14 open-weight models (VLMs, CUAs, VLAs) on 5,630 itch.io and 104 PyWeek games, finding that models make meaningful progress on only a minority of games, with Qwen3-Omni attaining the highest mean Progress Score without history on both corpora (0.84 and 1.26) and reaching at least Partial progress in 19.5% and 37.0% of runs, alongside recurring failures in spatial grounding, action execution, and self-correction.

Source-provided article image: PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence
Figure 2 ·

Figure 2 : Examples from the PlaySuite game corpus. PlaySuite contains a broad collection of independently developed games with varied visual styles, mechanics, camera perspectives, and controls. These examples illustrate the range of settings agents must handle, from 2D platform and puzzle games to racing, exploration, shooting, and 3D navigation environments.

arXiv

Interpretation

It constructs an interactive visual intelligence benchmark far larger than prior game benchmarks: 5,734 open-source games from PyWeek and itch.io, spanning engines such as Pygame, HTML5, Godot, and Unity across nine genres, with these independently developed games largely out-of-distribution for current models. Prior game benchmarks span only single digits to roughly two dozen games (for example Balrog with 6, Orak with 12, VideoGameBench with 23); PlaySuite is over two orders of magnitude larger and covers VLM, CUA, and VLA model families at once. The paper documents the full pipeline from an initial pool of about 1,460 PyWeek and 1.7M itch.io games through rating gates, platform compatibility, safety triage, and launch validation down to 5,734 games, and states that the final evaluation uses 5,630 itch.io and 104 PyWeek games.

It introduces a milestone-based Video-LLM-as-a-judge protocol that maps observable gameplay events to five ordered progress levels from None to Completed, without environment-specific rewards. Earlier multi-game evaluations relied on per-title engineering and hand-crafted rewards; this protocol uses fixed rules to map score increases, entry into new areas, intermediate objective completion, and game completion into a common progress scale, enabling comparison across thousands of heterogeneous games. Using Qwen3.5-9B as the default judge, validation on 60 gameplay videos (45 itch.io, 15 PyWeek) against six human annotators gives 93.0% human-human agreement within one level versus 90.4% for the judge; Qwen3.5-27B reaches 91.6%, while Gemma-4-31B-it and Molmo2-8B show lower agreement.

It provides empirical results across 14 open models showing that static visual reasoning and computer-use competence do not reliably transfer to goal-directed interaction in games. Prior capability demonstrations concentrated on a small number of environments; this work compares VLMs, CUAs, and VLAs under one protocol and reports that performance does not separate cleanly by family: CUAs remain competitive with similarly sized VLMs, while the evaluated VLAs generally trail the strongest VLM and CUA models. Without history, Qwen3-Omni attains the highest mean Progress Score on itch.io and PyWeek (0.84 and 1.26), with at least Partial progress in 19.5% and 37.0% of runs; with interaction history, Gemma-4-26B leads on both corpora (1.02 and 1.37). By genre, Simulation and Strategy score highest while Adventure and Puzzle score lowest.

It supplies HPC-oriented batched closed-loop execution infrastructure and several controlled analyses covering history length, prompting strategy, action frequency, and benchmark scale. A shared vLLM inference server with many concurrent game workers amortizes model loading across games; sampling experiments quantify how benchmark size relates to the stability of model rankings. On a single H100 with Qwen3-VL-8B, eight concurrent workers raise throughput from 6.7 to 42.4 evaluations per GPU-hour, and launch overhead drops from about 90 s to 8.3 s per game; median rank correlation with the full-corpus ranking is 0.846 at 50 games and 0.987 at 2,000 games, while the full-corpus top-three set is recovered in only 45.5% of 2,000-game samples.

Perspective

The benchmark targets zero-shot inference, measuring out-of-the-box generalization rather than maximum performance after fine-tuning or reinforcement learning, and is meant for researchers and engineering teams who want to compare general interactive models under one protocol. Evaluation runs on 5,630 itch.io and 104 PyWeek games, at 0.33 Hz for itch.io and at 0.33 Hz and 3 Hz for PyWeek, with each itch.io run spanning about 300 seconds of game time and roughly 100 interaction steps and PyWeek runs capped at 10,000 frames. The closed-loop protocol suspends game progression during model inference, so results reflect decision quality rather than real-time reaction speed. Milestone scoring captures observable progress and supports large-scale screening and trend comparison; the authors also note the pipeline currently validates approximately 17.5K executable environments, leaving room for future expansion.

Several open questions remain for a careful reader. Judge-human agreement within one level is 90.4%, below the 93.0% human-human rate, indicating that the five-level progress labels still leave room for judgment in borderline cases; variation across judge models also suggests some sensitivity to judge choice. The mean Progress Score summarizes ordinal outcomes, and the authors explicitly state it should not be interpreted as the fraction of a game completed. The counterintuitive finding that more history lowers average progress is accompanied by the observation that individual models can still benefit from history, so how temporal context can be used effectively remains unresolved. Finally, because the benchmark uses publicly available games, source code, and metadata, data contamination cannot be ruled out; incorporating newly released games can reduce but not eliminate this risk.

Sources