AeroEval stages program-level and execution-grounded validation to raise AI-generated drone mission success from 11/20 to 19/20
Synopsis
AeroEval is an agent-assisted middleware that validates AI-generated drone missions in stages between generation and execution: deterministic syntax and platform-API checks plus an intent-validation agent screen the program, accepted candidates run in a physics-based simulator, and a numerical trajectory comparator or a context-grounded agent then evaluates obstacle avoidance, altitude, coverage, orientation, and event-driven transitions, returning structured failure diagnoses for iterative regeneration; it raises navigation success from 11 to 19 of 20 tasks and aggregate analytical run-level success from 0.30 to 0.90.
((a))
arXivInterpretation
A staged validation lifecycle separates program validity, mission-intent alignment, and realized physical behavior, returning stage-specific feedback for iterative regeneration. Prior drone code-generation systems rely mainly on prompt guardrails or simulator outcomes with limited failure localization; AeroEval explicitly separates failures from program structure, platform API usage, mission interpretation, and realized physical behavior. Each validator returns a record of stage, outcome, violation, evidence, and correction; a failed record is supplied to the generator with the original candidate bundle, and the regenerated program re-enters Program Validation.
Context-grounded behavioral validation is developed for analytical missions without a unique reference trajectory, with validation agents combining mission requirements, program artifacts, robot and world descriptions, and execution observations to evaluate spatio-temporal and event-driven properties. Reference-path comparison only checks whether one expected path is reproduced; this method checks whether observed behavior satisfies mission properties, supporting inspection, survey, and search-and-track missions that admit multiple valid trajectories. The agent receives the mission specification, generated program, an ordered list of odometry tuples, robot constraints such as maximum flight speed and altitude, payload capacity, and battery limitations, and world constraints including obstacles, road network, permitted flying corridors, no-fly regions, and maximum allowed height; the current implementation infers properties directly from the textual world description and execution trace without separate geometric helpers.
Stage complementarity is measured: Code Validator-only and Trajectory Validator-only configurations achieve mean run-level success rates of 0.50 and 0.60, while their composition reaches 0.90. Program-level and execution-grounded checks capture different failure classes, so the composition gain comes from complementarity rather than one stage being stronger. Compared under a common budget of at most three regeneration cycles over 10 independent runs for each of five analytical mission types; 12 Delivery, 27 Farm Survey, and 64 Radio Tower Inspection candidates pass the Code Validator but fail the Trajectory Validator, quantifying code-valid but behaviorally invalid candidates.
End-to-end reliability improves and replicates across two reasoning models: navigation success rises from 11/20 to 19/20, and aggregate analytical run-level success rises from 0.30 to 0.90 with o3-mini; with Gemini-3.1-Pro, AeroGen succeeds in 9 runs and AeroEval in 35 runs, an increase of 52 percentage points. Relative to the closed-loop CLG baseline and one-shot AeroGen, staged diagnosis converts most previously unsuccessful tasks into accepted missions at the cost of more refinement iterations on average. Navigation tasks use AirSim with a 15-iteration budget and analytical tasks use Gazebo with a 15-iteration budget; AeroEval succeeds in 45 runs with o3-mini and 35 runs with Gemini-3.1-Pro, and because the confidence intervals overlap the results support staged validation across both model families rather than establishing that one model is more reliable.
Perspective
The work targets settings that insert a validation layer between AI mission generation and drone execution, suited to teams that already have platform API metadata, robot descriptions, and an executable physics-based simulator, such as drone application developers, cyber-physical validation researchers, and edge-analytics deployers. In the evaluated environment the method raises navigation success from 11/20 to 19/20 and aggregate analytical run-level success from 0.30 to 0.90, and supervised runs on a DJI Ryze Tello in an open circular field of approximately 50 m radius show that accepted Delivery, Farm Survey, and Search and Track programs can transfer to the evaluated platform. Supporting another programming language requires only a parser and platform-specific API metadata, leaving the validation-agent interfaces unchanged. An optional deployment advisor recommends analytics placements or waypoint schedules under constructed resource contention but does not determine mission acceptance.
The Trajectory Validator produces three false acceptances and no false rejections against 15 expert-labelled outcomes comprising 12 expert-rejected and 3 expert-accepted cases, and the authors state that the Context-Grounded Agent cannot serve as a standalone assurance mechanism; the labelled set is small and imbalanced, so only raw cross-model agreement is reported rather than a chance-corrected coefficient. Each analytical mission type uses only 10 runs, so per-mission intervals remain wide. Generation and validation default to the same model, which may retain correlated errors. Physical experiments use one platform and a limited set of mission and environment conditions, and trajectory extents and travel distances are scaled down because of the Tello's limited battery life, so they do not establish general sim-to-real transfer or safety. Validation overhead and cost vary by mission and model, and a converged Radio Tower Inspection run can take hours end to end. Readers who need exact per-mission success rates, confidence-interval values, and cost tables should consult the original figures and tables, since a fast parse may omit those details.
