Skip to main content
Back to timeline
arXivSource publication:

Uber's Persona Guardrail lifts agent out-of-domain detection from 57.3% to 93.5% with an iterated allowlist, cutting false approvals to 4.7%

Related research and updates

Synopsis

Uber researchers present Persona Guardrail, a production runtime defense framework that enforces each agent's functional boundary through synchronous input and output validation, together with PAGE, a benchmark of 2,076 labeled messages covering user and agent turns across four AgentDojo domains; on the non-mixed PAGE subset, an iteratively refined allowlist raises overall accuracy from 85.7% for a generic LLM judge to 95.9%, lifts out-of-domain detection from 57.3% to 93.5%, and cuts the false-approved rate from 25.0% to 4.7%, while the benign false-blocked rate stays essentially flat at about 3.6%–3.7%, and the framework runs in production within a sub-100 ms P90 per-call latency budget.

Source-provided article image: Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems
Figure 1 ·

Figure 1: Guardrail-call latency versus offered load on a sin-

arXiv · Page 10

Interpretation

The framework centers on a functional boundary: each agent's permitted and prohibited intents are written as semantic allowlist and blocklist specifications, and every request is validated before reaching the agent and every response before returning to the user. Where existing runtime defenses target specific attack classes such as prompt injection, this work reframes the defense objective as whether the agent stays within its intended functionality, covering both malicious requests and benign but out-of-scope requests. The paper gives a formal definition partitioning the query space into a should-process set and a should-reject set, and describes how the allowlist is generated by a code-oriented LLM analyzing the agent implementation and product requirements, then supplemented from production traffic in shadow mode.

On PAGE, the iterated-allowlist configuration achieves the strongest aggregate results: 95.9% accuracy, 96.3% precision, 95.3% recall, 95.8% F1, a 4.7% false-approved rate, and a 3.6% false-blocked rate. Relative to the baseline judge given only a generic agent description (85.7% accuracy, 57.3% out-of-domain detection, 25.0% false-approved rate), the explicit allowlist raises out-of-domain detection to 83.3%, and one refinement pass driven by inspecting false-approved and false-blocked turns raises it further to 93.5%. All configurations are evaluated on identical PAGE splits, and the appendix reports held-out validation on disjoint two-way and three-way partitions where accuracy rises monotonically across v1, v2, and v3 on unseen data (for example 91.9% to 96.1% to 96.7%), indicating the gains are not tuning overfitting.

A blocklist-only configuration is markedly weaker: 71.2% accuracy, only 14.5% out-of-domain detection, and a 49.0% false-approved rate; adding a blocklist on top of the allowlist does not help either, roughly doubling the benign false-blocked rate from 3.6% to 8.2% and slipping accuracy to 94.9%. This indicates a blocklist enumerates attacks rather than the agent's function, so it cannot recognize fluent, task-like requests that are merely out of scope, which is precisely the hard part of functional-boundary enforcement. Per-label analysis shows adversarial detection is high across allowlist configurations (96.8%–97.0%), while out-of-domain detection separates the configurations: 57.3% baseline, 83.3% allowlist, 93.5% iterated allowlist, and 14.5% blocklist.

For production deployment, the framework uses Qwen3-30B-A3B (a mixture-of-experts model with roughly 30B total and about 3B active parameters per token, quantized to FP8) as a shared classifier for both input and output guardrails, served with vLLM on NVIDIA Dynamo with a disaggregated CPU frontend and GPU workers. The paper treats the guardrail as a specialized inference workload, single-token and prefill-dominated with a static shared prefix accounting for about 93% of input tokens, and meets the latency budget through prefix-cache reuse, frontend/decode disaggregation, and bounded contexts rather than generic serving tuning. Between 5 and 50 QPS, median latency ranges from 19.43 to 25.34 ms and P90 latency from 57.30 to 61.08 ms, with P99 crossing the 100 ms objective at 30 QPS; accelerator utilization stays below 65%, and the disaggregated design reduces guardrail-call latency by approximately 30% relative to the prior general-purpose serving stack.

Perspective

The framework targets the threat model of direct malicious-user and out-of-domain interactions: an external user with black-box access through the public interface who cannot manipulate the agent's internal tools, retrieved documents, memory, or execution environment. In that setting it suits customer-facing agents that should reject out-of-scope requests, as in the four domains used here (Workspace, Slack, Banking, Travel); whether out-of-domain requests are blocked is configurable, so deployments that answer or gracefully redirect them are also supported. Methodologically, the allowlist is generated by a code-oriented LLM analyzing the agent implementation and product requirements, then supplemented from production traffic in shadow mode and refined through a human-readable intent taxonomy rather than classifier retraining; candidate updates must satisfy predefined security, utility, reliability, and latency thresholds across golden, adversarial, and production-sampled held-out pools, pass human review, ship as declarative configuration, and be evaluated in shadow mode before a limited enforcing rollout. Engineering results apply to single-accelerator model calls with a 4,096-token context cap, and the latency measurements exclude application and network overhead.

The paper states that its results are a defense-in-depth control for the defined threat model and do not establish complete protection against indirect prompt injection, compromised components, unauthorized intermediate tool actions, multi-turn attacks, or other threats outside that model, and that deployment decisions and acceptable error rates remain application-specific. The labeling semantics for mixed-intent turns is an open question: the guardrail blocks about 92.2% of mixed turns under both policies, so scoring them as should-approve inflates the false-blocked rate to 17.1% and drops accuracy to 88.6%, whereas scoring them as should-block keeps accuracy at 95.6% with a 3.6% false-blocked rate, meaning whether a turn carrying legitimate content should be approved needs to be settled per deployment. The informed-attacker experiments in the appendix show verbatim intent-list access gives attackers the greatest advantage, though that scenario carries relatively low real-world risk, and the closed-source attacker models refused 17%–71% of first-turn attacks due to safety tuning, a measurement condition that shapes how attack strength should be read. The exact production guardrail prompt cannot be released, so exact reproduction depends on internal infrastructure, and future work points to conversation-level and tool-call-level enforcement, neither of which is yet evaluated.

Sources