Skip to main content
Back to timeline
arXivSource publication:

RAG-Stress diagnostic protocol: instructing models to prioritize documents makes them replace correct answers with evidence-supported foils, with Strict exceeding Soft by 10.9–13.5 percentage points

Synopsis

RAG-Stress introduces a controlled diagnostic protocol that holds the question and reference answer fixed, edits one assertion so the evidence supports a designated incorrect answer, and crosses two source-priority policies with three positions of the answer span in the evidence text, measuring a misleading rate on questions each model answered correctly without retrieval across fifteen systems, three QA datasets, and English and Chinese medical QA. Instructions that prioritize documents consistently produce higher misleading rates than those permitting prior knowledge, with a mean gap of 10.9–13.5 percentage points across the three QA datasets, and model-averaged rates follow End, Beginning, Middle under both policies.

Source-provided article image: RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
Figure 1 ·

Figure 1: One saved trajectory: the same edit, two instructions, different answers. With clean evidence, both soft and strict prompting return the reference (Tomb Raider); with the edited evidence, soft prompting keeps the reference while strict prompting adopts the foil (Uncharted). The closed-book answer is also Tomb Raider.

arXiv

Interpretation

The work measures adoption of a designated foil conditional on closed-book correctness, defining a misleading rate (MR) that separates answer replacement from preexisting error. Prior evaluations ask whether retrieval improves accuracy or whether an answer is supported by retrieved text; aggregate accuracy collapses replacement with preexisting failure. RAG-Stress fixes the question and reference, edits one answer-bearing assertion to support a same-type foil, and scores on the baseline-correct subset. The protocol runs on fifteen systems (seven API systems, four open-source instruction models, four RL-trained search agents) over TriviaQA-RC, HotpotQA, and SearchQA, plus 1,000 sampled questions per language in English and Chinese MedQA; clean accuracy uses the full prepared set while MR uses each model's baseline-correct subset, and the two denominators are never subtracted.

Source-priority policy systematically shifts the misleading rate: Strict, which makes documents the primary source of truth, exceeds Soft, which permits prior knowledge to override conflicting documents, in every reported cell of the main tables. Within each comparison the question, evidence, and placement are fixed and only the source-priority policy changes, so the gap is consistent with evidence following that depends on the policy rather than different retrieved content; the direction is expected from the instructions, and the diagnostic result is the frequency of designated-foil adoption among baseline-correct answers. The gap averaged over models and positions ranges from 10.9 to 13.5 percentage points across the three QA datasets; on English and Chinese MedQA, Strict MR exceeds Soft MR for all 15 systems in both languages, with similar group-averaged gaps (API 12.8 vs. 12.4; open-source 16.9 vs. 18.3; RL-based 9.8 vs. 9.3).

The position of the answer span within the evidence text covaries with the misleading rate: averaged equally across the fifteen models, both instructions yield End, Beginning, Middle on all three datasets. Position effects in long contexts previously motivated but did not directly test word placement within a document; this work makes Word Position a crossed condition and adds a paired 300-question position check. Strict MR by position is 37.0/31.4/30.3% on TriviaQA-RC, 41.3/35.8/32.8% on HotpotQA, and 35.0/29.3/26.9% on SearchQA; in the paired check, Strict End-minus-Middle is 10.9 and 9.6 percentage points for models A and B, while all Beginning-minus-Middle and Soft contrast intervals include zero, supporting higher End MR under Strict rather than the complete ordering or a tested instruction–position interaction.

A separate paired audit of 500 questions, two quantized 8B checkpoints, and 13,000 saved generations supports increased harmful override under Strict without establishing a corresponding gain in beneficial correction. The audit reports harmful override (HOR), beneficial correction (BCR), and clean-to-edited transitions separately, with uncertainty estimated over questions while retaining all repeated cells, which distinguishes compliance with a source-priority policy from factual reliability. Relative to Soft, Strict increases foil adoption on the baseline-correct subsets by 14.0 and 9.7 percentage points, with HOR contrast intervals excluding zero for both checkpoints while the corresponding BCR intervals include zero; clean evidence produces almost no designated-foil matches (Strict 0.0%, Soft 0.0–0.3%) whereas the edited arm produces many more.

Perspective

The protocol is meant for readers who need to evaluate how retrieval-augmented systems behave when evidence conflicts with prior knowledge, under settings where evidence is fixed, the question and reference are fixed, and a closed-book baseline-correct stratum can be defined; the authors state that it measures a behavioral transition, not stable parametric knowledge or, by itself, the causal effect of editing. The main intervention varies the answer span's position within the target evidence text with document order fixed, whereas the separate audit moves the entire target passage to document slot 1, 2, or 4, so its position contrasts describe document order rather than word placement within the text. English and Chinese MedQA contain different questions rather than paired translations, so language differences covary with closed-book coverage and remain descriptive. The authors also note that the work does not identify an internal attention mechanism or a superior mitigation.

Several open questions remain for a careful reader: the fifteen-system main-table contrasts are descriptive, paired bootstrap intervals apply only to the 300-question QA checks and the 500-item audit, and they are not multiplicity-adjusted; the position check supports higher End MR under Strict but not the complete ordering or an instruction–position interaction; higher HotpotQA MR is consistent with a multi-hop reasoning-chain mechanism, yet evidence style, difficulty, and baseline-correct cohorts also differ, so matched bridge-versus-terminal edits would be needed to test it; HOR intervals excluding zero while BCR intervals include zero is not a test of their difference and does not establish that the BCR effect is exactly zero; the audit's human-review fields for semantic validity remain unfilled and its source is a historically Llama-selected 1,000-item pool, limiting generalization; and the decoding-profile check is a descriptive consistency check that does not isolate a sampling or output-length effect.

Sources