Public articles linked to the same research event.
arXiv The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.
The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.
The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.
The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.