WAND estimates wind disturbance with a TCN to drive both policy conditioning and feedforward compensation, raising quadrotor success by 8.3 percentage points on average across 12 wind-disturbed simulation settings and succeeding in 18 of 20 indoor fan tests
Synopsis
The paper proposes WAND, a reinforcement learning framework for quadrotor navigation in dense obstacle fields under time-varying wind: a Temporal Convolutional Network estimates wind-induced disturbance acceleration from proprioceptive histories, the estimate is injected into the policy through a zero-initialized residual module (WindAdapter) and reused for low-level feedforward compensation, improving observed success rate by 8.3 percentage points on average over feedforward compensation alone across 12 wind-disturbed simulation settings and succeeding in 18 of 20 indoor fan-induced flight trials.
Fig. 1 : Real-world deployment of WAND. (a) A quadrotor navigates a cluttered indoor scene under fan-induced airflow; the inset shows an anemometer reading of approximately 7.2 m / s 7.2\,\mathrm{m/s} . (b) The corresponding visualization in RViz.
arXivInterpretation
WAND uses the wind-disturbance estimate both to condition the navigation policy and to provide low-level feedforward compensation, so navigation actions depend explicitly on the estimated disturbance rather than inferring time-varying aerodynamic effects implicitly from obstacle and proprioceptive observations alone. Prior learning-based navigation policies condition mainly on obstacle perception and proprioception; WAND grounds temporal conditioning in a supervised estimate of wind-induced disturbance acceleration, making the conditioning signal physically interpretable and fusing it with obstacle features and the current proprioceptive state. The method defines the disturbance acceleration (Eq. 6) and a causal TCN with three residual causal blocks, specified kernel size and dilations, and left padding ensuring no access to future measurements, validated in simulation and real flight.
The TCN disturbance estimator achieves lower RMSE than EKF and LSTM under all four wind conditions, including a held-out mixed-wind dataset not used for training, validation, or model selection. Earlier wind or external-force estimation relied on proprioceptive or visual-inertial measurements and adaptive control; here a causal TCN is trained on a single supervisory target (simulator-applied aerodynamic force divided by mass), transferred from Gazebo to Isaac Sim without retraining, and frozen during PPO optimization. Table I reports RMSE for Constant, Turbulent, Gust, and Mixed wind; TCN values are 0.0777, 0.1876, 0.1159, and 0.1970, all below EKF and LSTM, with mixed wind evaluated on a separately generated held-out composite set.
Ablations show the zero-initialized WindAdapter and the directional-clearance reward term are complementary: neither zero initialization alone nor random initialization markedly improves over Base+Comp, while combining zero initialization with the reward term clearly raises the late-stage reach-goal rate. The reward term couples the estimated disturbance direction with LiDAR directional clearance so the safety margin adapts to disturbance magnitude and direction, while zero initialization lets the wind-conditioned residual be introduced progressively during joint training. Six configurations share identical environment randomization and PPO settings, comparing No Wind, Base, Base+Comp, w/o r_c, Random Init, and WAND training curves (Fig. 4).
In a fixed obstacle layout under opposite crosswinds, WAND succeeded in all 20 trials per wind direction, whereas Base+Comp achieved 60% and 0% success, with trajectories shifting to the opposite side of the central obstacle cluster when the crosswind reversed. This provides evidence of disturbance-conditioned navigation beyond feedforward compensation: the policy not only compensates the disturbance but changes its route and safety margin according to wind direction. Table II reports success/collision/timeout rates and minimum obstacle clearance with 20 trials per condition; under one wind direction WAND also increased the mean minimum clearance of successful trials.
Perspective
The work targets micro quadrotor navigation in dense obstacle fields under time-varying wind, applicable to spatially uniform constant, turbulent, gust, and mixed wind profiles in simulation and to controlled indoor fan-induced airflow; the policy transfers directly from simulation, the estimator is fine-tuned with about 10 minutes of real flight data, and the full model has roughly 342,698 parameters running onboard at 30 Hz, indicating a target setting of real-time onboard deployment on SWaP-constrained platforms. For readers, it offers a reference design for feeding an explicit disturbance estimate into both the policy and low-level control, useful for autonomous flight tasks that need wind-adaptive safety margins.
Wind profiles are spatially uniform in simulation and do not model obstacle-induced wakes or recirculation, and the authors list systematic evaluation under spatially varying wind fields as future work; real experiments use indoor fan-induced airflow with ten trials per scene, limiting disturbance intensity and spatial structure. The estimator captures only the disturbance at the vehicle's current location, no explicit delay compensation is applied in real deployment, and the acceleration output is used without numerical differentiation or additional filtering. In addition, some equations, reward weights, and clearance values appear as placeholders in the text, so precise reproduction should consult the original figures and equations.
