Public articles linked to the same research event.
arXiv The work builds PlaySuite, a benchmark curated from PyWeek and itch.io spanning 5,734 open-source games, together with a unified closed-loop interaction framework, HPC-oriented batched execution, and a Video-LLM-as-a-judge milestone scoring protocol, and evaluates 14 open-weight models (VLMs, CUAs, VLAs) on 5,630 itch.io and 104 PyWeek games, finding that models make meaningful progress on only a minority of games, with Qwen3-Omni attaining the highest mean Progress Score without history on both corpora (0.84 and 1.26) and reaching at least Partial progress in 19.5% and 37.0% of runs, alongside recurring failures in spatial grounding, action execution, and self-correction.
The work builds PlaySuite, a benchmark curated from PyWeek and itch.io spanning 5,734 open-source games, together with a unified closed-loop interaction framework, HPC-oriented batched execution, and a Video-LLM-as-a-judge milestone scoring protocol, and evaluates 14 open-weight models (VLMs, CUAs, VLAs) on 5,630 itch.io and 104 PyWeek games, finding that models make meaningful progress on only a minority of games, with Qwen3-Omni attaining the highest mean Progress Score without history on both corpora (0.84 and 1.26) and reaching at least Partial progress in 19.5% and 37.0% of runs, alongside recurring failures in spatial grounding, action execution, and self-correction.
The work builds PlaySuite, a benchmark curated from PyWeek and itch.io spanning 5,734 open-source games, together with a unified closed-loop interaction framework, HPC-oriented batched execution, and a Video-LLM-as-a-judge milestone scoring protocol, and evaluates 14 open-weight models (VLMs, CUAs, VLAs) on 5,630 itch.io and 104 PyWeek games, finding that models make meaningful progress on only a minority of games, with Qwen3-Omni attaining the highest mean Progress Score without history on both corpora (0.84 and 1.26) and reaching at least Partial progress in 19.5% and 37.0% of runs, alongside recurring failures in spatial grounding, action execution, and self-correction.
The work builds PlaySuite, a benchmark curated from PyWeek and itch.io spanning 5,734 open-source games, together with a unified closed-loop interaction framework, HPC-oriented batched execution, and a Video-LLM-as-a-judge milestone scoring protocol, and evaluates 14 open-weight models (VLMs, CUAs, VLAs) on 5,630 itch.io and 104 PyWeek games, finding that models make meaningful progress on only a minority of games, with Qwen3-Omni attaining the highest mean Progress Score without history on both corpora (0.84 and 1.26) and reaching at least Partial progress in 19.5% and 37.0% of runs, alongside recurring failures in spatial grounding, action execution, and self-correction.