EscapeGuard cuts tool-selection hallucination by 9.0 pp across six benchmarks and six models while suppressing cross-configuration hallucination escape
Related research and updatesSynopsis
The work identifies and formally defines hallucination escape, in which existing tool-hallucination mitigations reduce hallucination on the tool configuration they are tuned on but increase it on others; it finds that models hold intrinsic tool-use tendencies whose conflict with the runtime configuration sharply raises hallucination and that existing methods reinforce these tendencies; it then proposes EscapeGuard, a training-free inference-time method combining conflict-aware gating with configuration-derived attention enhancement, which reduces tool-selection hallucination by 9.0 pp, lowers the cross-configuration mean by 23.7 pp, and achieves 89.1% net improvement in paired-query evaluation across six benchmarks and multiple models.
Figure 1: Hallucination escape and repair. (a) Under the standard configuration ( C 1 C_{1} ), an existing mitigation method produces the correct call. (b) After the tool is renamed ( C 2 C_{2} ), the same method still calls the original name. (c) EscapeGuard adapts to the renamed configuration and produces the correct call.
arXivInterpretation
It identifies and formally defines hallucination escape, a previously overlooked failure mode in which a mitigation reduces hallucination on its tuned configuration but increases it on others by enough to outweigh the reduction. Prior work evaluated improvements only under the single tool configuration each method was tuned on, without accounting for other configurations; this work evaluates five existing methods along both tool-configuration variation and query variation and finds all of them satisfy the escape definition. Five methods (Relign, PALADIN, Gorilla, LinSteer, PRISMS) are reproduced on LLaMA-3.1-8B-Instruct and Qwen-3.5-9B using roughly 1,000 queries from BFCL V3 and Seal-Tools across six configurations and three random seeds; the methods reduce hallucination by 7.5 pp on average under their tuned configuration but collectively increase it by 13.0 pp across the remaining five, leaving every method with a higher cross-configuration mean than the base model.
It reveals the mechanism of hallucination escape: models hold stable intrinsic tool-use tendencies even without the runtime tool configuration, hallucination rises monotonically with tendency strength when those tendencies conflict with the current configuration, and existing methods reinforce these tendencies. It moves the explanation of hallucination from out-of-distribution configuration novelty to a conflict between intrinsic tendencies and the runtime configuration, and validates that conflict as separable in representation space via layerwise probes. Closed-book elicitation shows mean selection priors of 0.61–0.78 and usage priors of 0.57–0.71; under conflict, hallucination rises monotonically across prior-strength quintiles, most strongly for tool selection and argument names; layerwise logistic-regression probes detect conflict with 85%–87% accuracy (selection) and 77%–79% (argument name) and predict hallucination with AUROC 0.81–0.84 and 0.74–0.75; existing methods increase prior strength by 0.04–0.14, with escape ratios up to 2.42.
It proposes EscapeGuard, a training-free inference-time method that redirects attention to the runtime tool configuration when a tendency conflict is detected, mitigating tool hallucination while preventing hallucination escape. Unlike training-based or fixed-direction interventions, EscapeGuard derives its corrective signal from the current tool configuration rather than a fixed tendency, so it adapts naturally to tool renaming, replacement, and schema changes and leaves the intrinsic prior unchanged. Across six benchmarks and six backbones, tool-selection hallucination falls by 4.3–14.7 pp (9.0 pp overall) and argument-name hallucination by up to 67% relative; the cross-configuration mean drops by 23.7 pp with variance falling about 40%, and paired-query evaluation shows 87.8%–90.4% net improvement; inference overhead is 4%–6% of base latency.
Perspective
The results target single-turn tool-calling scenarios with external tools, in runtime settings where tool names, descriptions, and argument schemas may change across deployments; EscapeGuard is gated by frozen conflict probes and intervenes only when the conflict score exceeds a threshold, making it suited to engineering teams that want more reliable tool calls without retraining models. It is validated on six benchmarks (BFCL, Seal-Tools, RelyToolBench, API-Bank, MetaTool, APIBench) and six backbones, with cross-configuration evaluation covering standard, tool replacement, description change, argument renaming, mixed change, and distractor injection, and it reports repair rate, transfer damage rate, and net improvement on paired queries so readers can judge applicability by how much their own tool configurations vary.
The tendency-conflict mechanism accounts weakly for argument-value hallucination, and EscapeGuard's improvements on argument-value hallucination are correspondingly modest, with residual overall hallucination dominated by argument-value errors; gains are smallest under distractor injection, which the authors attribute to confusable tools diluting the enhanced attention signal. The probes retain predictive power under a novelty-matched control, but the authors note that fully ruling out all forms of distributional shift as a confound would require intervening on tendency strength while holding the configuration fixed, which is not straightforward in a pretrained model. In addition, the six tool configurations and tested architectures, though representative, are not exhaustive, and addressing argument-value errors, improving distractor robustness, and extending to broader architectures remain open directions.
