CrashSim generates crash scenarios from real-world crash priors, builds the 4,000+ scenario nuCrash dataset, and widens measured safety gaps across five planners
Related research and updatesSynopsis
The authors present CrashSim, a human behavior-informed crash scenario generation framework that uses a vision-language model to interpret naturalistic driving scenes, hierarchically retrieves real-world crash priors, and converts their relative-motion patterns and behavior-stage timings into anchor guidance for a diffusion generator pretrained on nuScenes; results show CrashSim reproduces real-world pre-impact behavior, collision dynamics, collision geometry and crash-type distributions more closely than Strive and LD-scene, and it is used to build nuCrash, a dataset of over 4,000 crash and near-crash scenarios on which five planners show larger closed-loop performance differences than on nuScenes, with an LLM evaluation agent providing capability-level diagnoses and improvement guidance.
Figure 1: Challenges in autonomous vehicle safety evaluation and overview of CrashSim. a Limitations of current autonomous vehicle safety evaluation, including the sparsity of real-world crashes and the behavioral, kinematic, and distributional biases of existing safety-critical simulation. Statistics in the real-world testing funnel are derived from ref. [ 32 ] . b Real-world crash processes evolve from intention formation through risk awareness and emergency response to the eventual collision outcome whereas existing safety-critical simulation often overlooks the emergency response and leads to aggressive collisions. c CrashSim retrieves real-world crash priors based on the current scene context and uses them to provide human behavior-informed guidance for generative multi-agent simulation and autonomous vehicle safety evaluation.
arXivInterpretation
CrashSim models real-world crashes as a temporally evolving process with four stages—intention formation, risk awareness, emergency response and collision outcome—rather than treating collision occurrence as a single generation objective. Existing environment-based, optimization-based, generative and LLM-based methods largely prioritize increasing collision occurrence or scenario criticality and give limited attention to the pre-impact interaction process, which can yield persistent aggressive motion, missing emergency responses or implausible vehicle kinematics. The paper compares behavioral-stage temporal composition, the temporal evolution of risk awareness and emergency response, pre-impact longitudinal acceleration, reaction time, response time and braking severity against real-world distributions, reporting that CrashSim agrees more closely than LD-scene and Strive.
CrashSim turns sparse, heterogeneous crash data into transferable behavioral guidance through hierarchical retrieval of real-world crash priors, generating more realistic safety-critical interactions in naturalistic scenes. Real-world crash data are sparse, unevenly distributed across interaction types and heterogeneous in representation, making a general-purpose crash generator hard to train; CrashSim instead builds a small-scale crash retrieval library and uses retrieval-augmented generation at inference time rather than a large, uniformly structured crash trajectory training set. Methodologically, the retrieval library is built from bird's-eye-view trajectory reconstructions of SHRP2 NDS crashes, assigned to 16 interaction modes derived from the NHTSA pre-crash typology and organized into mode, event and sample layers; retrieval is a three-stage cascade using all-MiniLM-L6-v2 sentence embeddings and cosine similarity, retaining a prior only when the composite sample-level similarity exceeds a threshold.
CrashSim also preserves real-world crash-type distributions at the scenario-set level and is used to construct the nuCrash long-tail dataset. Existing methods rarely account explicitly for the scene-level distribution of generated crashes and can over-represent particular conflict types; nuCrash increases the prevalence of safety-critical events while retaining real-world behavioral and distributional characteristics. The paper measures distributional discrepancy with Total Variation distance and Jensen-Shannon divergence over sixteen fine-grained crash types and five high-level conflict categories, reporting the smallest discrepancy for CrashSim; nuCrash contains over 4,000 crash and near-crash scenarios, has a higher crash rate than nuScenes and SimEngine, and shows smaller TV distance and JSD than SimEngine.
In closed-loop evaluation nuCrash separates planner safety capabilities more clearly than nuScenes, and an LLM evaluation agent converts deterministic metrics and scenario evidence into capability diagnoses and improvement directions. Conventional autonomous-driving evaluation frameworks rely primarily on predefined quantitative metrics, whereas this work offers a capability-oriented interpretation linking performance changes to concrete behavioral patterns and scenario conditions. The paper reports the normalized performance range rising from 0.08 on nuScenes to 0.58 on nuCrash and between-planner variance rising from 0.001 to 0.065; learning-based planners mainly show increased collision severity while rule-based planners show more pronounced degradation in motion stability or route completion; ablations show removing crash-prior retrieval reduces the emergency-response proportion from 22% to 12% (real-world value 32%), and removing the VLM agent raises TV distance from 0.093 to 0.281 and JSD from 0.078 to 0.185.
Perspective
The work targets simulation-based scenario generation and closed-loop planner evaluation for autonomous-vehicle safety, in a setting that starts from nuScenes naturalistic scenes and uses SHRP2 NDS reconstructed crash trajectories plus the NHTSA pre-crash typology as priors. Its two outputs have broader utility: the retrieval database built from real-world crash data can supply empirical crash priors to other diffusion-based or LLM-based scenario generation methods, and nuCrash extends naturalistic driving evaluation to safety-critical conditions rarely represented in conventional datasets while preserving the scene and trajectory representation used for nuScenes, enabling planner comparison under the same closed-loop protocol. For researchers and engineering teams who need to examine planner capability differences under long-tail hazardous interactions, nuCrash offers a more discriminative evaluation setting than naturalistic driving data alone; the LLM evaluation agent organizes deterministic metrics, failure patterns and scenario contexts into six-dimensional capability profiles and points to improvement directions and subsequent validation.
The LLM evaluation agent's attribution of failure mechanisms remains relatively coarse, and the paper itself lists deeper attribution analysis as future work; implementation details such as the specific prompts, normalization and weighting coefficients for crash-prior retrieval and VLM scene understanding are placed in supplementary material, so the full pipeline is hard to reproduce from the main text alone. In addition, nuCrash generation starts from nuScenes scenes, so its crash-type distribution is influenced by the relatively large number of adjacent-lane interactions in the source scenes, for example a comparatively high proportion of lane-change scenarios; using nuCrash as a planner training resource, and incorporating broader crash databases and more reactive surrounding-traffic modeling, remain open questions.
