Skip to main content
Back to timeline
arXivSource publication:

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

Synopsis

Using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across Gemma-3 and Qwen3.5 on six harmful-content benchmarks plus Spanish and Hindi-English code-mixed evaluations, this work separates failures of representation from failures of routing, finding that sparse readouts outperform native prediction on all six binary tasks (Qwen 0.740 vs 0.432 native macro-F1; Gemma 0.532 to 0.714), that probe-discriminative and output-routed directions dissociate, that calibration-only routing recovers 93.

Source-provided article image: Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Figure 1 ·

Figure 1: Residual-SAE construction and readout pathways. A. A frozen base SAE reconstructs layer- ℓ \ell states, while a trainable SAE models the normalized reconstruction residual, yielding base and joint reconstructions. B. The resulting representation supports native scoring, residual-reconstruction hooks, sparse probing, calibrated output routing, and probe-distilled LoRA. Snowflakes and flames denote frozen and trainable components, respectively.

arXiv

Interpretation

Sparse readouts outperform native prediction on all six binary tasks in both Gemma-3 and Qwen3.5, the most informative token role varies by task, and fine-grained labels remain substantially harder. Prior harmful-meme work typically optimizes the final detector without disentangling representational from readout failures; this work compares native constrained prediction with sparse probes on the same frozen model, making decodability an independently measurable quantity. Locked-split macro-F1 comparisons: Gemma mean 0.532 to 0.714, Qwen mean 0.432 to 0.740; all six Qwen probes use the same public layer-20 SAE, and Gemma replicates the pattern under a distinct model and sparse-representation regime.

In Qwen, probe-discriminative and output-routed directions dissociate under intervention: silent-feature ablation changes the probe 24-63 times more than the monitored yes/no margin, while routed-feature patching changes the native margin 16-140 times more than the probe. Prior probing work often equates decodability with model use; this work uses independently selected output-aligned features and causal interventions to split discriminative from routed directions, noting that the region of both high probe importance and high direct output alignment is empty on five of six analyzed tasks. Ablation and donor-target patching on Qwen's public base SAE, with 15 eligible examples per task for silent-feature knockout and 5 donor-target pairs per task for routed-feature patching; the authors state these ratios depend on the two score scales and are not scale-invariant mediation fractions.

This separation is actionable: calibration-only routing raises Qwen mean macro-F1 from 0.4405 to 0.7204, close to the probe ceiling of 0.7404, recovering 93.3% of the mean gap, and probe-distilled LoRA improves native prediction although a shared seven-task adapter causes negative transfer. Prior work largely stops at diagnosing internal-external discrepancies; this work additionally tests two recovery paths, output-level direct routing and internalizing the readout into model parameters, and compares dedicated versus shared adapters. Routing improves all six binary tasks on matched Qwen reporting subsets, with four tasks selecting the largest tested coefficient beta=8; a dedicated FHM adapter raises test-seen macro-F1 from 0.6482 to 0.7138, while the shared adapter reaches 0.6548 versus 0.7095 for the dedicated one on the same development set and reduces six-class MMHS150K from 0.3727 to 0.1886.

Gemma-3-12B shows a model-specific distributed rank-32 image-prompt interaction on FHM, reaching 0.756 versus 0.685 native, and cross-lingual, code-mixed, OCR-ablation, and image-perturbation controls indicate the effect extends beyond English, is not explained solely by supplied OCR, and depends materially on paired visual evidence. FHM's benign confounders defeat single-role and simple pooled interaction readouts; this work explicitly models image-position by prompt/OCR-position interaction with a low-rank bilinear term and uses leave-one-factor interventions to show the interaction is distributed rather than attributable to one feature pair. The calibration-selected rank-32 checkpoint reaches 0.7560 versus 0.6853 native on the reported test, with a paired bootstrap 95% confidence interval of [0.0461, 0.0954]; 11 of 32 factors have measurable leave-one-out effects but removing any single factor changes validation macro-F1 by at most 0.007; image permutation reduces probe macro-F1 by 0.148-0.206 and routed macro-F1 by 0.160-0.292.

Perspective

The work targets research and engineering settings that use frozen vision-language models for harmful meme classification; its readout-gap definition is limited to the analyzed layers, SAE dictionaries, features, and output anchors, and its recovery paths are validated on matched Qwen reporting subsets and a Gemma-3-4B-IT distillation setting. Multilingual conclusions rest on fixed internal holdouts, and the FHM interaction claim is limited to the analyzed Gemma-3-12B layer-31 representation. It enables next steps such as testing larger models, additional layers and dictionaries, broader multilingual and fine-grained datasets, and exploring position-matched patching and path-level tracing.

A careful reader would still watch whether the probe-versus-native comparison, which is not like-for-like, translates into deployable moderation reliability; whether the mechanistic conclusions hold beyond the selected layers, dictionaries, features, and output anchors; that some interventions use only 12-15 eligible examples and the ratios depend on score scales; that token roles identify position provenance rather than modality-pure states and generated-token probes are post-generation readouts; that the FHM interaction does not transfer to every model; and that multilingual results rely on internal holdouts while rare fine-grained classes remain poorly resolved. In addition, although the full text was loaded, some figures and appendix details are presented in summarized form, so verifying exact sample counts and per-task effects would still require the original appendices.

Sources