Skip to main content
Back to timeline
arXivSource publication:

Sentry retrieves failure lessons on demand at test time, beating the strongest runtime-intervention baseline by 37% on average across agentic benchmarks

Related research and updates

Synopsis

The work introduces Sentry, a failure-management layer running alongside an LLM agent: when it detects a failure it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did, while the full playbook never enters the agent's context; across multiple agentic benchmarks Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated, combining Sentry with context evolution yields further gains, learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance.

Source-provided article image: Sentry: Learning to Recover from LLM Agent Failures at Test Time
Figure 1 ·

Figure 1: Two ways failure knowledge goes wrong (Qwen3.5-9B on WebShop). (a) Always exposed: recovery rules learned by ACE stay in the agent’s context; after finding a matching item, the agent keeps comparing alternatives and takes no action. (b) Never learned: AgentGuard’s fixed no-progress warning resolves a search loop, but no lesson is stored; in a later task, the same warning fires after valid option selections, and the agent leaves the page and loses them.

arXiv

Interpretation

The paper argues that failure knowledge is conditional knowledge and should be conditionally exposed: a lesson should enter the agent's context only when its corresponding failure occurs. Prior approaches either keep failure lessons permanently in context (context evolution) or intervene only when a failure occurs without learning from the repair (runtime intervention); this work makes the timing of exposure itself a design principle. Controlled experiments show that exposing the full playbook to the agent lowers performance even when relevant lessons remain available on demand, and that removing lessons from an evolving playbook improves performance.

Sentry is a failure-management layer running alongside the agent that closes the loop of detecting a failure, retrieving matching lessons, verifying recovery without rewards, and storing a new lesson only on success. Runtime interventions repair but do not learn, and context evolution learns but keeps knowledge resident in context; Sentry separates learning from exposure, and the full playbook never enters the agent's context. Across multiple agentic benchmarks, Sentry outperforms the strongest runtime-intervention baseline on every benchmark by 37% on average, and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated.

Lessons learned by Sentry transfer to held-out tasks, and combining Sentry with context evolution yields further gains. This indicates that on-demand retrieval of failure lessons is not overfitting to specific tasks but a complementary mechanism that can be layered onto existing context-evolution methods. The paper reports transfer results on held-out tasks and further gains when Sentry is combined with context evolution.

Perspective

The result targets LLM agent systems that need improved reliability at test time, in settings where detectable failure signals exist and an external playbook can be maintained; the method is designed as a failure-management layer running alongside the agent, so its benefits presuppose that the agent can be observed and intervened upon. For practitioners who want to add recovery capability without modifying the agent itself, and for readers studying context management and test-time adaptation, the conditional-exposure principle is directly applicable; the further gains when combined with context evolution indicate the two mechanisms can coexist.

The abstract does not give the specific benchmark names, task scale, distribution of failure types, or statistical significance, nor the concrete criterion used to verify recovery; the size of the playbook, the retrieval matching method, and the granularity of lesson writing are likewise not detailed in the abstract. Readers who care about these details will need the main text's benchmark setup and ablations. In addition, the exact composition of the 'multiple agentic benchmarks' and how held-out tasks are split are open questions for judging the scope of the transfer claim.

Sources