Public articles linked to the same research event.
arXiv Uber researchers present Persona Guardrail, a production runtime defense framework that enforces each agent's functional boundary through synchronous input and output validation, together with PAGE, a benchmark of 2,076 labeled messages covering user and agent turns across four AgentDojo domains; on the non-mixed PAGE subset, an iteratively refined allowlist raises overall accuracy from 85.7% for a generic LLM judge to 95.9%, lifts out-of-domain detection from 57.3% to 93.5%, and cuts the false-approved rate from 25.0% to 4.7%, while the benign false-blocked rate stays essentially flat at about 3.6%–3.7%, and the framework runs in production within a sub-100 ms P90 per-call latency budget.
Uber researchers present Persona Guardrail, a production runtime defense framework that enforces each agent's functional boundary through synchronous input and output validation, together with PAGE, a benchmark of 2,076 labeled messages covering user and agent turns across four AgentDojo domains; on the non-mixed PAGE subset, an iteratively refined allowlist raises overall accuracy from 85.7% for a generic LLM judge to 95.9%, lifts out-of-domain detection from 57.3% to 93.5%, and cuts the false-approved rate from 25.0% to 4.7%, while the benign false-blocked rate stays essentially flat at about 3.6%–3.7%, and the framework runs in production within a sub-100 ms P90 per-call latency budget.
Uber researchers present Persona Guardrail, a production runtime defense framework that enforces each agent's functional boundary through synchronous input and output validation, together with PAGE, a benchmark of 2,076 labeled messages covering user and agent turns across four AgentDojo domains; on the non-mixed PAGE subset, an iteratively refined allowlist raises overall accuracy from 85.7% for a generic LLM judge to 95.9%, lifts out-of-domain detection from 57.3% to 93.5%, and cuts the false-approved rate from 25.0% to 4.7%, while the benign false-blocked rate stays essentially flat at about 3.6%–3.7%, and the framework runs in production within a sub-100 ms P90 per-call latency budget.
Uber researchers present Persona Guardrail, a production runtime defense framework that enforces each agent's functional boundary through synchronous input and output validation, together with PAGE, a benchmark of 2,076 labeled messages covering user and agent turns across four AgentDojo domains; on the non-mixed PAGE subset, an iteratively refined allowlist raises overall accuracy from 85.7% for a generic LLM judge to 95.9%, lifts out-of-domain detection from 57.3% to 93.5%, and cuts the false-approved rate from 25.0% to 4.7%, while the benign false-blocked rate stays essentially flat at about 3.6%–3.7%, and the framework runs in production within a sub-100 ms P90 per-call latency budget.