PLUTO turns natural-language rendezvous missions into SCP constraints, generating a trajectory for all 50 prompts with 94% passing mission-intent verification
Related research and updatesSynopsis
The work introduces PLUTO, an interactive trajectory design framework that embeds agentic AI coding tools within a sequential convex programming (SCP) architecture, mapping natural-language rendezvous mission requirements into structured mathematical constraints that pass through an auto-convexification pipeline into executable optimal control formulations; across 50 natural-language rendezvous mission prompts spanning 10 constraint types, PLUTO generated a trajectory for every prompt and 94% of the resulting trajectories satisfied both numerical constraint checks and semantic consistency with the original mission intent.
Fig. 1 : PLUTO (Plain Language Understanding for Trajectory Optimization) architecture. The agent parses a natural-language prompt into convexity-classified constraint objects, which the optimization layer sorts, convexifies, and solves via SCP. Results are displayed in the dashboard for user review and iteration.
arXivInterpretation
PLUTO integrates requirement interpretation, optimal control problem formulation, numerical solution, and verification into a unified interactive workflow, so users describe desired spacecraft behavior in natural language without supplying equations, constraint classes, or solver code. Prior LLM trajectory work largely operates over fixed optimization formulations, with task-specific constraints predefined by the system designer or not explicitly incorporated into the optimization problem; PLUTO lets operator intent modify the formulation at the constraint level rather than only selecting among predefined configurations. The paper describes this workflow through four layers (user, agent, optimization, and visualization/verification) with structured interfaces, and details how constraint synthesis, the ConstraintSet, and verification artifacts are organized.
An auto-convexification pipeline linearizes differentiable non-convex constraints to first order about the current SCP reference point, using JAX for automatic differentiation and CVXPY to build convex subproblems, with re-linearization as SCP iterations progress. This design lets natural-language-synthesized constraint functions be injected directly into the SCP solve: convex constraints enter the subproblem directly, while supported non-convex constraints enter after linearization, extending expressible constraints while retaining the computational structure of convex optimization. The paper gives the formulation of a non-convex constraint and its first-order approximation, and lists the mathematical forms and convexity classes of 10 constraint types (T01–T10), with T01 spherical/ellipsoidal keep-out zones, T07 minimum maneuver magnitude, and T10 range-dependent speed, glide slope, and out-of-plane exclusion among the non-convex ones.
Across 50 prompts (30 single-constraint, 10 two-constraint, 10 three-constraint), solver completion was 100%, constraint formulation matched the reference in 90%, optimal convergence was 84%, and prompt verification was 94%. The evaluation reports both intermediate-stage and end-to-end metrics, showing that formulation deviations or non-optimal termination did not translate proportionally into prompt verification failures, indicating robustness to imperfections in intermediate layers. A reference formulation was manually constructed and solved for each prompt, required to be feasible, produce a trajectory, and satisfy the same numerical and semantic verification criteria; all experiments used the same initial condition and time horizon with claude-sonnet-5 as the agentic model.
Failure modes concentrated in the T10 class of complex or uncommon constraints: all three prompt-verification failures exhibited both a formulation mismatch and non-optimal convergence and involved T10, with two stemming from the same out-of-plane exclusion constraint where the agent encoded the full disjunction instead of a one-sided convex half-space simplification. The paper attributes such failures to the engineering-judgment layer—choosing the simplest formulation sufficient to meet a requirement rather than over-engineering—and notes this kind of 'tools of the trade' knowledge can be encoded into PLUTO's reference materials. The paper reports that 2 of 5 formulation-deviation prompts still passed verification, and that 6 of 8 non-optimal terminations came from prompts containing T10; all prompts ran on a clean directory with no stored context, keeping evaluation conditions equal.
Perspective
The framework targets spacecraft engineers and mission designers who describe requirements in natural language and want to rapidly construct and inspect rendezvous trajectories; it applies to rendezvous scenarios built on a relative orbital elements state, impulsive maneuvers, and an SCP backend. Its value lies in letting operators modify the formulation at the constraint level while retaining explicit numerical verification and visual review. The paper notes that PLUTO's context-rich structure can encode 'tools of the trade' simplifications into its reference materials, so such failures are expected to become less frequent as lessons from prior runs accumulate, potentially yielding more efficient constraint representations; future work will investigate incorporating lessons from prior runs toward recursive self-improvement and evaluating model-agnostic capability with open-source agents.
The evaluation fixes claude-sonnet-5 as the agentic model, so model dependence (including open-weight models) remains unassessed and conclusions apply to this framework-model combination. Test scenarios share one initial condition, time horizon, and baseline keep-out constraint, with a 30/10/10 prompt distribution, so extrapolation to other dynamics, mission types, or constraint libraries still needs validation. Constraint formulation is scored against manually constructed references, and natural language admits multiple valid interpretations, so that metric reflects interpretive consistency rather than correctness. The paper notes that visual outputs and similar results are harder to capture numerically, so readers may want to consult the dashboard and constraint-geometry plots when judging semantic verification. In addition, the failure analysis rests on a small number of cases (three verification failures, eight non-optimal terminations), so its characterization of failure mechanisms is best read as directional.
