Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Sentry retrieves failure lessons on demand at test time, beating the strongest runtime-intervention baseline by 37% on average across agentic benchmarks

The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.