Skip to main content
Back to timeline
arXivSource publication:

AdaGuard uses an adaptive guard model to review agent trajectories against user-defined policies, with its 4B model reaching 89.30% binary accuracy on AdaptiveSafety

Synopsis

The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.

AI-generated editorial illustration: AdaGuard: An Adaptive Guard Model with User-defined Policies

Interpretation

It introduces AdaptiveSafety, a dataset that links changes in policy and in recorded behavior to the corresponding violation judgments. Whereas existing guard datasets annotate prompts and responses under predefined safety criteria, this dataset combines policy counterfactuals, which change permissions, conditions, thresholds, or exceptions while keeping the trajectory fixed, with behavioral counterfactuals, which keep the policy fixed while changing recorded facts such as authorization, refusal versus execution, and tool outcomes, and adds structural augmentations that supervise consistency under rule reordering and identifier remapping. The dataset contains 10,939 training examples and 1,000 test examples, with policies ranging from 1 to 100 rules; each example pairs the policy and trajectory with an overall analysis and the complete list of violated rule identifiers in policy order, using NR to indicate no violation; the test set is excluded from supervised training, reinforcement learning, and checkpoint selection.

It proposes SafePO, a reinforcement learning algorithm that refines violation identification through structured rewards and value-guided token weighting. Because a binary compliance reward cannot distinguish a complete verdict from one that detects a violation but omits other applicable rules, SafePO evaluates the predicted violation set and its serialization separately, gives full credit to an exact verdict, distinguishes ordering errors from incorrect rule predictions, and scores partial matches using set overlap and sequence agreement, while penalizing malformed responses and identifiers outside the supplied policy. For each policy and interaction the algorithm samples a group of responses and centers their rewards at the group mean and scales them by the group standard deviation to obtain response-level advantages; an independently trained value model adjusts token weights within the analysis and verdict regions while keeping each region's total weight fixed, giving the short verdict a prescribed share of the learning signal; it uses clipped policy updates and KL regularization toward the frozen supervised model.

It develops AdaGuard models at 0.6B, 4B, and 8B parameters that take a user-supplied policy at inference time and produce an overall analysis followed by a verdict identifying violated rules. Compared with prior adaptive guards whose primary focus remains conversational compliance and content safety, this work extends assessment to tool-mediated agent behavior, requiring joint interpretation of the supplied policy and the trajectory and returning the complete set of violated rules rather than only a binary judgment. The 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench; the 8B model reaches 89.50% accuracy and 88.67% binary F1 on AdaptiveSafety, identifies the complete rule set in 77.10% of examples with 74.05% rule micro-F1, and on DynaBench reaches 76.80% accuracy, 76.05% F1, and 70.72% rule exact match.

It reports stratified results by policy length and assessment scope, along with a comparison of training stages. Results are broken down by policy-length group and separated into user-only requests versus trajectories, and the saved SFT checkpoints are compared with the SafePO checkpoints, rather than reporting only a single aggregate score. With 51 to 100 rules, the 8B model reaches 88.40% accuracy and 77.20% exact match; within AdaptiveSafety the 117 user-only requests and 883 trajectories are evaluated separately, with the 4B model at 91.45% accuracy on requests and 89.01% on trajectories; from SFT to SafePO the 8B model improves AdaptiveSafety accuracy from 88.40% to 89.50% and DynaBench accuracy from 74.77% to 76.80%, while some metrics at 4B and 0.6B decrease.

Perspective

The work targets deployment settings where agent behavior must be assessed at inference time against user-defined policies, applicable when the policy is given as an ordered list of natural-language rules and the interaction record contains user requests, agent messages and actions, and observations returned by tools or the environment; the dataset and models cover policies with 1 to 100 rules and distinguish user-only requests from trajectories. For developers who want rule-level attribution rather than only a binary judgment, AdaGuard offers an overall analysis plus violated rule identifiers, and only the actor is required at inference time. The authors note that deployment should retain human review for consequential decisions and that performance should be assessed under the policies of the target application.

All reported evaluations are single runs, and the authors provide no significance claims or estimates of training-seed variability; the factual correctness of generated explanations is not independently verified, since the reward measures verdict correctness and output structure. The SFT-versus-SafePO comparison is descriptive evidence across saved checkpoints, the historical evaluator used a different input-budget calculation, and component ablations isolating the reward terms, value modulation, or region budgets are not available. On DynaBench, AdaGuard-8B remains below DynaGuard-8B in accuracy and F1, and the API models Jev and GPT-5.2 also score higher there, with different model interfaces and inference budgets across comparisons. Policy-length groups differ in content and label composition, and exact match is not monotonic in policy length. These scope notes and open questions indicate that which training choices drive the gains, and how reliably the resulting guards support intervention under a target application's policies, still await controlled experiments and broader deployment studies.

Sources