Skip to main content
Back to timeline
arXivSource publication:

DroneWAM cuts drone visual-navigation inference to 483 ms with JEPA latent prediction and adaptive rollout, lifting closed-loop progress from 0.510 to 0.774

Synopsis

The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.

AI-generated editorial illustration: DroneWAM: Efficient World Action Model for Drone Visual Navigation

Interpretation

DroneWAM predicts the future in representation space rather than generating future images, and compresses visual tokens with a pretrained Resampler, lowering both the spatial and temporal cost of prediction. Compared with aerial world models and WAMs that explicitly predict the future visual world, and with vision-language navigation methods that model future visual consequences only implicitly, efficiency is built directly into the predictive process. Component ablations show Rel.ATE dropping from 0.6621 to 0.0584 once the Planner is added; adding the Resampler cuts inference from 656.12 ms to 582.65 ms with almost unchanged trajectory error; compressing from 512 to 128 tokens lowers latency from 591.28 ms to 483.29 ms while Rel.ATE moves only from 0.0574 to 0.0578.

Adaptive rollout uses a preference-trained Gate to choose prediction depth per scene, improving trajectory accuracy while reducing computation. Existing models commonly use a predefined rollout horizon and assign similar predictive computation to different scenes; here the rollout becomes variable-length with a stopping decision. Average rollout falls from 8.0 to 4.578 while Rel.ATE and Rel.RPE improve over fixed depth; random stopping at nearly identical average depth 4.581 performs considerably worse, indicating the gain comes from allocation rather than truncation; per-sample analysis shows the optimal depth varies across samples.

DroneNav-6D supplies 6-DoF aerial motion supervision beyond the conventional 4-DoF abstraction. Most established aerial visual-navigation benchmarks represent flight with 4-DoF translation and yaw, giving limited coverage of roll and pitch, which directly affect camera viewpoint and visual dynamics. The corpus contains 3,837 trajectory clips, 1,592,170 RGB observations, 63.03 hours of flight, and 2,004.11 km of travel, split into 3,453 training and 384 validation clips with all five simulation worlds represented in both, plus randomized wind disturbances.

The benefit of adaptive rollout extends beyond aerial navigation to LIBERO robotic manipulation. The same rollout strategy is transferred to a different embodied task distribution to test whether it is specific to aerial scenes. On LIBERO, latency drops from 890.8 ms to 411.9 ms (53.8% reduction), normalized RMSE from 0.2685 to 0.2652, and gripper accuracy rises from 89.83% to 90.55%, with consistent gains across all four suites.

Perspective

The result targets drone visual navigation under constrained onboard computation, especially image-goal closed-loop flight; the authors report that on moderate-complexity easy routes DroneWAM reaches an 88% success rate, 14 and 16 percentage points above NWM and FastWAM, with task progress 0.95 at 460 ms latency. The applicability of adaptive rollout is further extended to robotic manipulation, suggesting that allocating predictive computation per input transfers to other embodied task distributions. The DroneNav-6D dataset is meant for training and evaluating aerial world-action models that need 6-DoF motion supervision, and the authors state that codes and data will be released.

In the limitations, the authors note that the fixed token budget and bounded imagination horizon may require adjustment across scenes and computational budgets, and that performance depends on the information retained by the pretrained encoder and resampler; closed-loop task success in complex simulated scenes remains very low, and the current system is not ready for real-world deployment. Per-sample case studies show the Gate occasionally stops too early and misses the lower error available at deeper rollout. On the challenging closed-loop routes none of the compared methods completes an episode, so these gains are better read as directional evidence than as established consensus. The readable text here is the full paper body and appendices, but tables appear as text, so exact formatting of individual values should be checked against the original.

Sources