DreamTest replaces the simulator with a world model: up to 97% higher failure-prediction AUPRC on Parking, Humanoid and DonkeyCar, and 79% more novel failures under the same budget
Synopsis
DreamTest adapts a recurrent state-space model into a testing world model that learns agent behaviour and environment dynamics from a DRL agent's training log, then scores a candidate configuration by generating imagined rollouts in parallel and combining termination-weighted failure predictions, guiding search without executing the simulator; on Parking, Humanoid and DonkeyCar its mean AUPRC exceeds the strongest baseline by about 97%, 12% and 39%, gains on five out-of-distribution sets reach 145%, 29% and 44%, and under the same simulator-validation budget the best DreamTest-plus-search combinations find 29%, 22% and 79% more novel failures on average.
Figure 1: Workflow of DreamTest .
arXivInterpretation
Introduces the first world-model surrogate for testing DRL agents: an RSSM is adapted to initialise episodes from configurations, behaviour-clone the tested agent's actions, and predict termination and failure, so that a query returns a failure score from imagined rollouts while keeping the configuration-to-score interface that existing search algorithms already use. Prior surrogates such as Indago's MLP and SAMOTA's Kriging, polynomial and RBF models are discriminative: they map configurations to pass/fail labels and discard the per-step interaction records in the training log. DreamTest learns agent and environment dynamics from those per-step records and rolls them forward at query time. Supported by an ablation that keeps the same trajectory-trained RSSM but replaces imagination with a linear failure classifier on the configuration representation: AUPRC falls from 0.156 to 0.121 on Parking and from 0.393 to 0.177 on DonkeyCar, both significant, while Humanoid shows 0.205 versus 0.210 with no significant difference, which the authors attribute to the Humanoid configuration being the agent's initial physical state.
Achieves the highest mean failure-prediction AUPRC in-distribution, exceeding the strongest baseline by roughly 97% on Parking, 12% on Humanoid and 39% on DonkeyCar, and leads on all five out-of-distribution sets, with 23 of 25 baseline comparisons significant; the two exceptions (Unseen slots and the Humanoid torso-5% tail) retain higher means without a significant difference. Prior evaluations of baseline surrogates treated labels recorded during training as ground truth; the paper identifies this staleness and re-executes held-out configurations with the final agent to obtain fresh labels, which change 2.1%, 7.8% and 10.1% of labels on Parking, Humanoid and DonkeyCar respectively. Ten stratified 80/20 training-test splits with a shared budget of 60 hyperparameter trials, reporting both historical and fresh labels; the Parking base rate drops to 1.4% under fresh labels and DreamTest still holds the highest mean.
Under an identical simulator-validation budget of 50 submitted configurations, the best DreamTest-plus-search combinations find 29% (Parking), 22% (Humanoid) and 79% (DonkeyCar) more novel failures on average than the strongest baseline combination; non-surrogate approaches are far weaker, with best means of 5.7, 4.0 and 0.7 versus 14.5, 21.5 and 12.2 for the best ensembled DreamTest. Prior evaluations counted every observed failure even though seeded searches start from logged failures and may regenerate known cases; this work counts only novel failures, defined as failing during validation with a configuration that does not match any logged failing configuration after rounding every coordinate to four decimal places, and reports replay frequency. Each experiment is repeated ten times, with both single models and ten-model ensembles; gains are significant on Parking and DonkeyCar but not on Humanoid, and training-failure replay yields zero novel failures by definition.
On failure diversity, sweeping the number of clusters from 2 to 40 shows DreamTest covering the most behavioural clusters for almost every value, so the advantage does not hinge on a single cluster count chosen by the silhouette coefficient; gains are largest and most consistent on Parking, clear on Humanoid except under hill climbing, and more search-algorithm-sensitive on DonkeyCar. Prior work typically reported coverage at a single cluster count, which can bias comparisons because results depend on partition granularity; this work uses identical cluster definitions for every technique and sweeps the full range. Failing executions are represented as time series sampled at every simulator step (ego position for Parking, torso height for Humanoid, cross-track error for DonkeyCar), zero-padded, flattened and clustered with k-means; under the DonkeyCar genetic algorithm DreamTest finds 3.8 novel failures and 1.1 clusters per run, with 33.0 of its 36.8 raw failures being exact replays, which the authors explain as a highly accurate surrogate trapping genetic search in a dense failure region.
Perspective
The result targets engineering and research settings that test trained DRL agents in simulation or on real systems, especially high-fidelity simulation, physical robot testing and autonomous-vehicle testing where each execution is costly. The method requires only the per-step records already produced during agent training (configuration, observations, actions, rewards, termination and pass/fail outcome), so it can be applied to other agents with such logs without additional simulator executions; test generation runs offline after training and adds no latency to the deployed agent. The setting it is meant for is one where failure depends on both the configuration and extended agent-environment interaction: Parking and DonkeyCar, whose outcomes depend on longer interaction, benefit most, whereas the Humanoid configuration is itself the initial physical state and imagination adds limited signal. The authors also note that the surrogate only needs to rank failure-prone configurations ahead of less informative ones, with the simulator remaining the oracle that confirms each outcome, so imagined trajectories need not match every state exactly.
Imagined-episode fidelity is limited: the action head reproduces the agent's actions with decreasing accuracy across Parking, Humanoid and DonkeyCar, lowest on DonkeyCar where the policy acts from camera images; after 30 steps the standardised RMSE of the mean imagined trajectory is 0.62, 0.48 and 0.55, lower than a persistence predictor and the mean training trajectory but still accumulating with the imagination horizon; on DonkeyCar the termination head can let an imagined car continue beyond the real crash, compressing failure scores toward zero. On Humanoid 93% of episodes hit the time limit, leaving little duration variation and a lower termination correlation. Search-algorithm and environment pairing matters: on DonkeyCar seeded searches spend much of the budget replaying known failures, and hill climbing takes 30-40 hours there because saliency-guided search repeatedly fails to construct a valid neighbouring road from failure seeds near the generator's validity boundary. Novelty and behavioural clusters also remain proxies for distinct faults and root causes, and single-execution labels and failure counts remain noisy; the study does not evaluate retraining with the discovered failure sets or validate generality beyond the three benchmarks. A reader with only a fast parse and no figures should check the original for the exact shape of the cluster-coverage curves and the crossings across search algorithms.
