Skip to main content
Back to timeline
arXivSource publication:

WorldAuditBench tests multimodal agents on 213 3D world-auditing tasks: best model 42.3% vs. humans 83.4%

Synopsis

The authors built WorldAuditBench, a benchmark of 213 tasks across 13 Unreal Engine 5 and Three.js environments and five anomaly families for interactive 3D world auditing, and compared a single-model VLM agent with a two-stage VLA–VLM auditor, finding VLM agents reach 28.2%–42.3% success versus 6.6%–17.4% for two-stage auditors, both far below the 83.4% human rate.

AI-generated editorial illustration: WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Interpretation

The paper introduces WorldAuditBench: 213 tasks across 13 interactive 3D environments (126 tasks in Unreal Engine 5, 87 in Three.js), each implanted with a single pre-defined, labeled anomaly, organized into five families—static physics, interactive physics, spatial consistency, temporal consistency, and semantic consistency—spanning fifteen categories. Prior benchmarks either evaluate visual reasoning over fixed observations (Video-MME, GlitchBench, VideoGameQA-Bench, VideoGlitchBench) or interactive task completion (SimWorld, UnrealZoo, EmbodiedBench, OpenEQA), while FlySearch covers anomaly search but mainly out-of-place objects; WorldAuditBench makes active evidence gathering itself the task, including anomalies exposed only through interaction, viewpoint changes, or revisiting a location over time. All tasks are hand-crafted by the authors: they start from a bug-free scene, define permissible exploration regions, the agent's starting position and orientation, and available interactions, then perturb the scene to introduce one anomaly. Each task carries a scene description, ground-truth anomaly category, and evaluation rubric. Human validation uses reviewers who did not construct the tasks, working through a browser interface; a task is accepted only after at least two distinct reviewers rate the same configuration as Pass, with Fail or Uncertain triggering revision and re-review.

Evaluating five frontier models under two auditing paradigms, VLM-only agents reach 28.2%–42.3% success while two-stage VLA–VLM auditors reach 6.6%–17.4%, against a human baseline of 83.4%. The paper treats whether action and visual reasoning are coupled as the experimental variable: the VLM-only agent can choose its next action as it reasons, whereas the two-stage paradigm first collects a fixed trajectory with a VLA and then analyzes it offline, so intermediate observations cannot guide further exploration. Coupled auditing performs better overall, suggesting that using observations to direct exploration helps agents gather evidence to verify suspected defects. The VLM-only paradigm receives a budget of 40 environment actions; the VLA–VLM paradigm uses Open-P2P 1.2B to collect 60 seconds of simulated exploration with observations sampled every 0.5 seconds, and recorded trajectories are shared across VLM backbones. Success is determined by GPT-6 Astra as a VLM judge against each task's rubric; the paper reports 91.72% overall agreement between the judge and human judgments (92.86% for Unreal Engine, 90.00% for Three.js). The human baseline involves approximately ten computer science PhD students, with each task evaluated by two participants under a ten-minute limit.

Ablations and decomposition analyses locate the gap: VLA exploration struggles to reach anomalies, while memory tools and task guidance substantially affect what VLM agents discover. The paper separates failure to expose an anomaly from failure to identify it once exposed, rather than reporting a single score. Geometric coverage is 91.1% for the VLM agent versus 47.4% for the VLA explorer, and human-rated exposure is 68.1% versus 34.7%. On tasks geometrically covered by both, the VLM agent reaches 38.5% versus 26.0%; on tasks with human-rated exposure, 62.1% versus 47.3%. Controlled ablations with Gemini 3.8 Flash on 126 Unreal tasks: moving the start closer raises VLM success from 33.3% to 38.1% and VLA–VLM from 5.6% to 7.1%; halving the budget drops VLM to 23.0% and VLA–VLM to 4.0%; removing the in-context example lowers VLM to 31.0%, and further removing the anomaly-type hint lowers it to 15.9% and VLA–VLM to 2.4%. In the multi-anomaly experiment, anchor success is 42.9%, 42.9%, and 47.6% for one, two, and three anomalies, but identifying every anomaly occurs in only 3/42 and 2/42 runs, with target recall of 29.8% and 29.4%.

Perspective

This work targets researchers studying intelligent behavior in interactive 3D worlds and practitioners in simulation or game QA; it applies where an agent explores first-person under a limited budget, interacts, and submits an anomaly report with supporting evidence. What it offers is an evaluation protocol and baselines: evaluation code, task configurations, and runnable environment packages are planned for release under applicable licenses, so follow-up work can compare new auditing strategies, memory mechanisms, or exploration planners under the same rubric. The paper also makes 'knowing what to look for' an actionable variable—the anomaly-type hint matters more than the in-context example—which gives direct guidance for designing task prompts and curricula.

The reported success rates come from specific model and harness combinations (each model paired with its own CLI and a shared MCP interface), so numbers may shift with other backbones or tool sets; the two-stage paradigm analyzes fixed recorded trajectories, so its performance partly depends on what those trajectories cover. The multi-anomaly experiment shows anchor success changing little with anomaly count while complete discovery stays rare, leaving the trade-offs of multi-target auditing under a fixed budget an open question. Temporal consistency anomalies stay at or below 12.5% success for every model and paradigm, and why this family is especially hard, and what memory or revisit mechanism would help, is not yet answered. In addition, some tables in the available text present model names and full per-family values in placeholder form, so exact per-model, per-family numbers should be checked against the original tables.

Sources