Skip to main content
Back to timeline
arXivSource publication:

Ctrl-CWM plans over imagined crowd futures with a world model, beating CrowdES on most crowd realism and collision metrics while enabling run-time avoidance and attraction control in simulation

Synopsis

The authors propose Ctrl-CWM, a multi-agent controllable crowd world model that first learns and freezes an encoder via trajectory prediction on real pedestrian videos to preserve human motion dynamics, then uses an actor to propose displacements, a critic to rank imagined crowd trajectories, and a planner combining critic scores with user costs to select actions, enabling crowd generation and run-time control without retraining; on ETH–UCY, SDD, and GCS it outperforms CrowdES on most crowd realism and collision metrics and achieves the highest average compliance in avoidance and attraction scenarios.

Source-provided article image: Controllable Crowd Generation through World-Model Planning
Figure 2 ·

Figure 2: Overview of Ctrl-CWM. Trajectory prediction trains the state encoder h θ h_{\theta} on real-world pedestrian data, which is then frozen. The actor π ψ \pi_{\psi} and critic V ϕ V_{\phi} are learned over imagined rollouts on this representation, and CEM adds a user cost to the critic score for run-time control.

arXiv

Interpretation

Ctrl-CWM transfers the world-model principle of planning over imagined futures to crowd simulation, using a single model for both crowd generation and run-time control. Prior learning-based methods either limit control to a predefined population and agent parameters or rely on gradient guidance and sample filtering; this work instead plans over imagined futures so that user costs can introduce new objectives during simulation without retraining. The method section defines the state, actions, actor, critic, and CEM planner, and compares against ORCA and CrowdES on ETH–UCY, SDD, and GCS with shared arrivals.

Pretraining via trajectory prediction and freezing the encoder transfers real-world human motion dynamics to subsequent behavior learning. Relative to a randomly initialized encoder, pretraining reduces kinematic error from 1.205 to 0.465 and increases diversity from 0.157 to 0.285; the prediction heads reach 0.21/0.31 m average ADE/FDE on ETH–UCY and 5.89/10.26 px on GCS under best-of-20 evaluation. The ablation replaces the pretrained encoder while retaining the same behavior-learning procedure and separately evaluates the prediction heads; the authors note this comparison alone does not establish that freezing is preferable to fine-tuning.

Critic-guided planning over imagined futures reduces collisions while preserving realism, with the user cost driving the control objective. Relative to actor-only generation, critic-guided planning reduces DTW from 0.970 to 0.925 and collisions from 0.931% to 0.756%, with a small increase in kinematic error from 0.432 to 0.465; the user-cost-only variant achieves higher compliance (0.858 vs. 0.824) but performs worse on all eight realism and accuracy metrics. Component ablations compare under the same avoidance protocol and post-activation evaluation window; removing collision or kinematic supervision lowers compliance to 0.645 and 0.772, respectively.

In run-time avoidance control, Ctrl-CWM achieves the highest average compliance across every tested zone configuration. For multiple disjoint zones, average compliance is 0.950 versus 0.579 for CrowdES and 0.087 for ORCA; it also leads for disc and rectangle zones. Ten trials per scene evaluated over the final two thirds, with identical arrivals across methods; the authors note baselines use obstacles from initialization, so intervention times differ and this is not a matched comparison of online objective activation.

Perspective

This work targets researchers and practitioners who need to change crowd objectives dynamically in simulation, such as robot navigation training, autonomous-driving scenario generation, and urban design evaluation. Results apply to settings where objectives are expressed as explicit spatial costs: avoidance zones and attraction targets enter planning through user costs, and changing the target or zone changes the planning objective without retraining the model or altering the static scene map. The authors report that default planning parameters (candidate count, CEM iterations, and planning horizon) balance compliance, latency, and DTW at about 10 ms per call, and provide weight and emitter-population sweeps showing that control strength and emission settings jointly affect the observed trade-off. The authors also note the current framework requires user objectives as explicit spatial costs rather than high-level semantic commands, with future work grounding language instructions into planning objectives.

Several open questions remain for a careful reader. Avoidance compliance at weights of at least 25 can reflect slowed or stopped motion rather than successful rerouting, as the authors explicitly note; stronger attraction raises kinematic error and collision rate substantially (mean changes of +7.76 and +26.5 percentage points at the highest setting), so high compliance should not be read as preserving realism or safety. The weight and population sweeps use one trial per configuration without variability estimates, and the authors describe their trends as descriptive rather than statistical significance claims. In the main avoidance comparison, baselines use obstacles from initialization, so intervention times differ from Ctrl-CWM's online objective activation and this is not a matched comparison. In addition, the pretraining comparison cannot determine whether freezing outperforms fine-tuning, and forecasting accuracy alone does not establish the accuracy of all counterfactual crowd responses encountered during control. Among generation metrics, Freq. and Cov. coincide on single-category data, so ranks across displayed metrics are not independent measurements.

Sources