Public articles linked to the same research event.
arXiv The work identifies and formally defines hallucination escape, in which existing tool-hallucination mitigations reduce hallucination on the tool configuration they are tuned on but increase it on others; it finds that models hold intrinsic tool-use tendencies whose conflict with the runtime configuration sharply raises hallucination and that existing methods reinforce these tendencies; it then proposes EscapeGuard, a training-free inference-time method combining conflict-aware gating with configuration-derived attention enhancement, which reduces tool-selection hallucination by 9.0 pp, lowers the cross-configuration mean by 23.7 pp, and achieves 89.1% net improvement in paired-query evaluation across six benchmarks and multiple models.
The work identifies and formally defines hallucination escape, in which existing tool-hallucination mitigations reduce hallucination on the tool configuration they are tuned on but increase it on others; it finds that models hold intrinsic tool-use tendencies whose conflict with the runtime configuration sharply raises hallucination and that existing methods reinforce these tendencies; it then proposes EscapeGuard, a training-free inference-time method combining conflict-aware gating with configuration-derived attention enhancement, which reduces tool-selection hallucination by 9.0 pp, lowers the cross-configuration mean by 23.7 pp, and achieves 89.1% net improvement in paired-query evaluation across six benchmarks and multiple models.
The work identifies and formally defines hallucination escape, in which existing tool-hallucination mitigations reduce hallucination on the tool configuration they are tuned on but increase it on others; it finds that models hold intrinsic tool-use tendencies whose conflict with the runtime configuration sharply raises hallucination and that existing methods reinforce these tendencies; it then proposes EscapeGuard, a training-free inference-time method combining conflict-aware gating with configuration-derived attention enhancement, which reduces tool-selection hallucination by 9.0 pp, lowers the cross-configuration mean by 23.7 pp, and achieves 89.1% net improvement in paired-query evaluation across six benchmarks and multiple models.
The work identifies and formally defines hallucination escape, in which existing tool-hallucination mitigations reduce hallucination on the tool configuration they are tuned on but increase it on others; it finds that models hold intrinsic tool-use tendencies whose conflict with the runtime configuration sharply raises hallucination and that existing methods reinforce these tendencies; it then proposes EscapeGuard, a training-free inference-time method combining conflict-aware gating with configuration-derived attention enhancement, which reduces tool-selection hallucination by 9.0 pp, lowers the cross-configuration mean by 23.7 pp, and achieves 89.1% net improvement in paired-query evaluation across six benchmarks and multiple models.