Reading answer-label logits under a chain-of-thought cue drops Qwen2.5-VL-7B on ScienceQA from 80.76% to 45.48%, with 93.54% of predictions landing in the first option slot
Synopsis
The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
Figure 1: A single suffix creates an event mismatch. Direct prompting requests and scores B; the CoT prefix requests a rationale, so immediate label scoring can return A even when free generation and a matched probe recover B.
arXivInterpretation
The paper identifies and names CoT-prefix scoring, an evaluation-interface mismatch in which a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. Prior discussion of chain-of-thought largely asked whether reasoning improves capability; this work locates the issue in the evaluation step where the requested output event and the scored output event diverge. Illustrated by the Qwen2.5-VL-7B comparison on ScienceQA, where accuracy falls from 80.76% to 45.48%, together with slot statistics across five option-content permutations.
Under this mismatch, predictions skew strongly toward the first option slot: across five option-content permutations, 93.54% of CoT-prefix predictions select the first slot. It links the accuracy drop to a position bias, indicating the decline is not random noise but systematic behavior tied to option ordering. Based on prediction-distribution statistics across five option-content permutations.
Answer information often survives the prefix: condition-matched linear probes recover 78.94% from the same hidden states, and free generation restores 75.24%. Two independent readouts show that failure of the immediate answer-label logit readout does not mean the model lacks the answer. Comparison of recovery rates from linear probes and free generation on the same hidden states.
Vocabulary and layer diagnostics offer a mechanistic account: probability mass moves toward continuation tokens while answer information remains linearly accessible in late layers; the effect recurs with varying severity across datasets and models, though not universally. It extends the phenomenon from a single model to cross-dataset and cross-model replication observations while explicitly noting it is not universal. Vocabulary-distribution and layer-probe diagnostics, plus repeated observations across datasets and models.
Perspective
This work speaks to researchers and benchmark maintainers who use chain-of-thought prompts for multiple-choice evaluation, in settings where the scorer reads answer-label logits before the model generates a rationale. Its actionable direction is to align the requested output event with the scored output event, or to use alternative readouts such as condition-matched probes and free generation when reading answers. The cross-dataset and cross-model recurrence suggests the check can extend to other multiple-choice VLM evaluation settings, but the text explicitly states the effect is not universal, so applicability should be judged per setting.
Readers should still watch which factors determine how severely the effect appears across datasets and models; under what conditions condition-matched linear probes and free generation come closer to true capability; and where scores stabilize once requested and scored output events are aligned. Because this is based on summary-level text, the specific dataset list, model list, probe setup, and statistical details are not expanded in the text, and these remain open questions that require the original paper.
