ICDP constrains interaction distribution shift via contrastive density-ratio estimation, suppressing high-value but interaction-unsupported trajectories on nuPlan closed-loop and a real truck
Synopsis
The work introduces ICDP, an offline reinforcement learning framework that exactly decomposes joint-support degradation over ego and surrounding-agent futures into an ego-support component and a residual interaction-support component, recovers the latter through contrastive density-ratio estimation, and constrains policy improvement without explicit joint-density modeling, surrounding-agent prediction, or reactive simulator or world-model rollouts; on nuPlan and InterPlan closed-loop evaluations and real-world truck experiments, ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios.
Figure 1: Overview of ICDP . The logged scene is encoded into structured scene features used by the diffusion actor and critic. The actor generates candidate ego trajectories, which are evaluated by a pessimistic chunk-level critic and by the interaction-support module. The latter combines a joint ego–agent classifier with an ego-only classifier; their residual score isolates interaction-specific support and yields the IDS estimate used to constrain policy improvement.
arXivInterpretation
The paper identifies and formalizes interaction distribution shift (IDS): in interactive offline RL, an ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent evolution observed in the logged interaction. Existing offline RL methods largely control distribution shift in the policy's own action space, through conservative value estimation, in-sample learning, or regularization toward the behavior distribution; this work argues such mechanisms do not characterize joint interaction support and isolates IDS as a distinct source of shift beyond marginal ego support. The identification rests on a conditional-factorization derivation: joint-support degradation is written exactly as an ego marginal-support term plus a residual interaction-support term, formally separating IDS from ego-trajectory shift.
The paper derives a contrastive density-ratio estimator whose residual interaction score exactly recovers IDS at the population optimum, and integrates this constraint with chunk-level value learning and diffusion-policy optimization in ICDP. Recovering IDS requires neither explicit high-dimensional joint trajectory density estimation, nor prediction of surrounding-agent futures, nor reactive simulator or learned world-model rollouts during policy optimization; the key is that both contrastive objectives share the same policy-generated candidate trajectories, so the marginal ego-shift component cancels by subtraction. The result is stated as a theorem: under balanced class sampling and common support, the difference of the population-optimal logits of the joint and ego classifiers, evaluated against the same logged surrounding-agent trajectory, exactly equals IDS; the appendix provides the proof and shows why sharing the candidate distribution is necessary. The expectation of IDS is further interpreted as the divergence between the conditional interaction distributions associated with the logged and candidate ego trajectories.
On the nuPlan closed-loop benchmark, ICDP achieves the strongest reported performance across Val14, Test14-Hard, and Test14-Random under both non-reactive and reactive evaluation, for example 78.34 non-reactive and 72.87 reactive on Test14-Hard, and 93.33 non-reactive and 86.34 reactive on Test14-Random. Compared on the same nuPlan benchmarks against representative planners including PDM-Open, GameFormer, PlanTF, PLUTO, Diffusion Planner, and Flow Planner, ICDP leads on all evaluations; gains persist under reactive evaluation, where surrounding agents respond to the ego vehicle and interaction quality becomes particularly important. Evidence comes from closed-loop benchmark tables on the same benchmarks, reporting variants without rule-based trajectory refinement where available to compare learned policies directly; under a controlled backbone and identical training data, the interaction constraint yields the strongest performance in five of six settings, further improving reactive scores on Test14-Hard and Test14-Random relative to unconstrained offline RL.
Ablations show that both ego-support and interaction-support regularization have a moderate-strength optimum, and that the interaction-support estimator benefits from a sufficiently broad set of surrounding agents with limited further gain beyond that; real-world truck experiments provide qualitative evidence of feasibility in interactive and nominal driving. These analyses connect the interaction constraint to regularization strength and multi-agent context size, and add deployment observations beyond simulation rather than a single benchmark score. Regularization sensitivity is reported as average closed-loop score across Test14-Hard and Test14-Random: ego-support regularization improves performance up to a point and then decreases as the weight grows further, and interaction regularization likewise performs better at a moderate value than at weaker or stronger settings; performance improves consistently as surrounding agents increase, with a small decrease when context is expanded further. Real-world experiments on a snow-covered test track cover multi-vehicle interactions (overtaking and lead-vehicle following) and standalone curved-road driving, which the authors position as qualitative evidence.
Perspective
The result targets an offline setting that improves only the ego policy from fixed driving logs, with surrounding traffic treated as logged environment evolution rather than predicted or jointly optimized; the interaction-support estimator uses a limited number of surrounding-agent trajectories in the final configuration, and regularization strength must be set at a moderate value balancing ego and interaction support. For a reader, this offers a route to constraining interaction-level distribution shift without simulator interaction or world-model rollouts, suited to safety-critical planning teams that hold interaction logs with surrounding-agent futures and want reward-driven improvement beyond demonstrations; the real-world truck experiments on a snow-covered test track cover multi-vehicle interactions and standalone curved-road driving, which the authors position as qualitative feasibility evidence. The authors propose extending ICDP to online reinforcement learning, refining behavior from closed-loop interaction while retaining interaction-aware support control.
A careful reader would still watch: exact recovery of IDS relies on conditions such as population optimality, balanced class sampling, and common support, and the gap between learned classifiers and these conditions affects how reliable the constraint is; how much surrounding-agent context the support estimator should use and which regularization weights to choose are described as empirically best at moderate values, and their stability across data distributions and scenario compositions remains an open question. The real-world portion is qualitative evidence, so quantitative on-road performance, behavior under different weather and sensing conditions, and the effect of support control after extending from offline to online reinforcement learning remain to be observed. In addition, several equations and table values are not fully rendered in the parsed text, so exact reproduction of experimental configurations and per-item scores should consult the original appendix and figures.
