Skip to main content
Back to timeline
arXivSource publication:

Constraint violations stop counting equally: an 8B planner hits 89.2% pass rate at 2.1s latency

Lead

Preference optimization now scales its training signal by how badly a constraint was violated: an 8B model reaches 89.2% hard-constraint pass rate on travel planning, close to a multi-agent system's 89.4%, while cutting latency from 28.5s to 2.1s.

Source-provided article image: CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
Figure 1 · arXiv

Story

The training signal now distinguishes how severe a violation is, so a $1 budget overshoot and a $1,000 overshoot no longer produce the same update. Direct preference optimization previously treated every rejected sample alike, giving the same implicit reward gap regardless of failure severity. On the 180 TravelPlanner queries, replacing the binary preference signal with a verifier-derived continuous severity raised pass rate from 64.1% to 72.6%.

Hard constraints and soft preferences are split into two levels, with hard constraints satisfied first and soft preferences optimized only inside the feasible region. Earlier methods mixed both kinds of constraint into one preference signal, which could trade a hard constraint away for a preference score. On TravelPlanner the model showed no errors where a soft preference outranked a hard constraint, and reached a soft-preference score of 0.64 among feasible plans versus 0.72 for the multi-agent system.

Training pairs come from a teacher that makes the minimum edit, changing only the entities responsible for the violation and leaving the rest untouched. Earlier distillation had the teacher rewrite the whole plan, so the student could pick up wording and length differences instead of the correction. Minimal edits cut the average change from 245 tokens to 32, and under standard preference optimization raised pass rate from 64.1% to 78.3%.

What to watch

The next step is to test whether this severity-weighted signal transfers to open-world settings where constraints cannot be checked deterministically, and whether it reproduces at substantially smaller or larger model scales. Engineers building planning agents can adopt the two-level objective directly in settings that have a symbolic verifier, keeping hard-constraint satisfaction ahead of soft preferences.

The verifier reliability audit reports its agreement, false-pass, and false-fail figures as placeholders in the loaded text, so the exact numbers need checking against the original. Budget violations account for the largest share of TravelPlanner failures, leaving implicit arithmetic in longer plans as a bottleneck to watch. Teacher-guided synthesis costs about $320 per 15K pairs, so whether larger-scale deployment pays off depends on more efficient self-play strategies.

Sources