Skip to main content
Back to timeline
arXivSource publication:

AgentWorld tests multi-agent collaboration on 100 long-horizon MMORPG tasks: the best model reaches only 52.0% success and a CCE of 0.320

Synopsis

The authors built AgentWorld, a benchmark of 100 human-annotated tasks (plus 100 LLM-augmented variants) on the open-source MMORPG engine Kaetram, requiring 3–20 agents with asymmetric roles to coordinate over 25–55 rounds under a blackbox setting, and proposed Causal Collaboration Effectiveness (CCE), a graph-based metric over causal action graphs; testing Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B, the best model reaches only 52.0% task success and a CCE of 0.320, with failure modes centered on communication breakdowns, role confusion, and inability to maintain shared plans across rounds.

AI-generated editorial illustration: AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Interpretation

It introduces AgentWorld, a long-horizon blackbox collaboration benchmark: 100 human-annotated tasks plus 100 LLM-augmented variants across 8 categories (combat, crafting, gathering, trading, exploration, survival, construction, coordination), spanning 25–55 rounds and requiring 3–20 agents with asymmetric roles. The paper argues existing benchmarks mostly test competitive settings, short horizons under 20 steps, or simply aggregate individual performance, failing to isolate genuine collaboration; AgentWorld wraps low-level mechanics (pathfinding, combat, harvesting) into 13 high-level API tools so scores reflect collaboration rather than action-control proficiency. Tasks were designed by five human annotators following a pre-defined category distribution, each with a Python verifier, and revised through pilot experiments with two LLMs until only agent errors remained; round budgets were calibrated from those pilot runs.

It proposes Causal Collaboration Effectiveness (CCE): a causal action graph is built over task trajectories, success actions are identified, and backward tracing with binary LLM causal judgments yields the fraction of all actions that contributed to the outcome. The paper notes success rate cannot distinguish a task where one agent does all the work while others idle, and that holistic LLM-as-judge scoring is subjective and shifts when the judge model is updated; CCE reduces the judge's role to objective binary causal decisions. On 84 action–contribution judgments, a human annotator agreed with the GPT-4.1 judge on 69 (82% raw agreement, Cohen's κ = 0.64); recomputing with Claude Sonnet 4 gave 84% action-level agreement over 7,298 decisions. Recomputing CCE with three judges (GPT-4.1, Claude Sonnet 4, Llama-4 Maverick) fully preserved the model ordering, with absolute values shifting by at most 0.05.

Results across four frontier models show collaboration remains a common gap: on the main set Gemini 3 Flash leads at 52.0% SR, followed by Claude Haiku 4.5 (45.0%), GPT-5 Mini (36.0%), and DeepSeek R1-70B (20.0%); the best model's CCE is only 0.320, meaning less than a third of its actions causally contributed. The paper reports that communication quantity does not predict success: GPT-5 Mini sends the most messages per task (44.1) but ranks third in SR, DeepSeek sends the fewest (7.5) and ranks last, while Gemini communicates moderately (11.0) yet achieves the highest SR; conditioning on successful tasks, Claude has the highest per-task collaboration efficiency (0.653 vs. Gemini's 0.609). All four models used identical system prompts, user prompt templates, and API tool definitions; 100 main tasks and 100 augmented variants were run; on the augmented set all models dropped substantially (Gemini 52%→24%, Claude 45%→26%, GPT-5 36%→21%, DeepSeek 20%→10%).

Baselines and ablations characterize where the difficulty lies: random actions solve 5.7%, the single-agent baseline 28.6%, no-communication 22.9%, shared-plan-no-communication is lower still (17.6%), and oracle communication raises success to 60.0%; removing task documentation (54.3%→29.4%) and shortening the round budget from 50 to 10 (54.3%→37.1%) cause the largest drops. The paper uses these controls to show the tasks genuinely require multiple coordinating agents and that the value of planning depends on being able to revise it during execution; even with oracle communication a 40% residual failure rate indicates much of the difficulty is intrinsic to long-horizon planning and resource allocation. Ablations ran on a 35-task subset; Gemini 3 Flash run three times at temperature 0.7 with seeds 42, 819, and 314 achieved 54.3% SR in each run (sample standard deviation 0.0 percentage points), with standard deviations of 2.07 rounds and 0.18 chats.

Perspective

The work targets researchers and engineering teams evaluating the collaboration capability of LLM agent teams, in settings that are blackbox, role-asymmetric, and long-horizon (25–55 rounds); the sandbox, task definitions with verifiers, evaluation scripts, and annotation platform are fully open-source and can be reused or extended. CCE adopts an intentionally inclusive labeling policy in which marginally helpful actions count as contributing, so the paper suggests reading CCE as an upper bound on the fraction of genuinely useful actions.

CCE still relies on an LLM for the underlying causal judgments; the paper supports this with human agreement and ordering stability across three judge models, but absolute values shift with the judge (e.g., Claude Haiku moves from 0.294 to 0.241 between GPT-4.1 and Claude Sonnet 4). The three-run repeat covers a single model on a 35-task subset, and the paper states this does not establish stability across models or statistical significance of model differences. Augmented variants were generated by Claude Opus, with human validation on a 10% sample, of which 90% were judged to have appropriate difficulty. The communication failure-mode analysis is based on 61 error messages from Gemini 3 Flash traces across 100 tasks and is qualitative. In addition, this reading is of the paper's full text, where figures appear as textual descriptions, so specific graphical details were not directly visible.

Sources