MSU team's pseudo-mask and local-contrastive system for the Explainable Deepfake Detection Challenge 2026 reaches 0.9349 test detection accuracy and a 0.7456 final challenge score
Synopsis
For the Explainable Deepfake Detection Challenge on XPlainVerse, the authors present a modular detection-and-explanation system in which a detector fuses several DINOv3 backbones with Mesorch forensic features, a Grounding-DINO pipeline converts local artifact descriptions from explanations into weak patch-level pseudo-masks supervising an Artifact Evidence Map, and a local patch-level contrastive objective separates artifact from authenticity evidence, while class-conditional Qwen3-VL models generate complex explanations and a GRPO-optimized model simplifies them, reaching 0.9349 detection accuracy, a 0.5571 explanation score, and a 0.7456 final challenge score on the full test split.
Figure 1. Overview of the proposed system, including offline pseudo-mask generation, multi-backbone forensic detection, and class-conditional explanation generation. A three-part pipeline: explanation-grounded pseudo-mask generation, a multi-backbone detector with image and patch losses, and class-conditional complex and simple explanation generation.
arXivInterpretation
A pseudo-mask pipeline turns local artifact cues in fake-image explanations into patch-level weak supervision: Qwen3-VL-32B-Instruct extracts short artifact phrases from complex explanations and builds a vocabulary split into localizable local attributes and image-wide global attributes, only local attributes are grounded with Grounding DINO, and the union of retained boxes is mapped onto the detector patch grid so each patch target is the covered fraction used to supervise the Artifact Evidence Map. Unlike detectors that rely on pixel-level manipulation annotations or image-level labels alone, this pipeline converts natural-language forensic cues into usable spatial supervision, and the authors state explicitly that the masks are not pixel-accurate manipulation annotations but regions where textual evidence and open-vocabulary grounding agree. In the public-validation ablation, the image-level detector alone reaches 0.8975 AUC, and adding pseudo-mask artifact supervision raises it to 0.9139, indicating a gain for the image-level decision.
A local patch-level contrastive objective (LCL) assigns each patch a consistent, uncertain, or opposite-side status from the image label and the current Artifact Evidence Map prediction, then uses bidirectional cosine matching across same-class and cross-class image pairs to select pull, push, or ignore, with stop-gradient preventing uncertain matches from moving both embeddings at once. The objective requires neither paired images nor pixel-level manipulation masks, moves contrastive learning from the image level to the patch level, and explicitly handles the uncertainty introduced by potentially incomplete pseudo-masks. In the ablation, adding LCL on top of pseudo-mask supervision raises validation AUC from 0.9139 to 0.9417, the largest single gain among the three loss configurations.
A multi-backbone forensic detector fuses three transformer-based DINOv3 models, three CNN-based DINOv3 models, and a Mesorch manipulation-localization backbone, aligning DINO spatial maps and the Mesorch spatial representation onto a shared patch grid as a Unified Forensic Feature Map, with max pooling letting localized evidence affect the image-level decision and a patch head producing the Artifact Evidence Map. Relative to single-backbone or purely image-level classifiers, the design brings together pretrained visual representations, DCT-aware cues, and multi-scale forensic information, using the frequency-aware, localization-oriented representation as a complement to DINOv3. On the final test split the submission reaches 0.9349 detection accuracy and 0.9340 macro F1, with fake and real F1 of 0.9418 and 0.9261.
Explanation generation uses separate class-conditional Qwen3-VL-8B-Instruct complex-explanation models routed by the detector prediction to the fake or real branch, followed by a text-only simplifier for general users that is optimized with GRPO against the official simple-explanation score after supervised fine-tuning. Rather than having one multimodal model both decide and explain, the system keeps classification in a specialized detector and language output in class-conditional generators to reduce class leakage, especially hallucinated artifacts on real images. On an internal held-out subset of 10K examples (5K real, 5K fake) using only generations with the correct final answer, the class-conditional generators raise entity F1 from the organizer baseline's 0.5536 to 0.6235 and evidence F1 from 0.4620 to 0.5447, while normalized SLE rises from 0.4264 to 0.9703.
Perspective
The results apply to the XPlainVerse challenge subset: 760,000 images with image-level labels and complex and simple explanations, on which both the detector and the explanation models were trained and evaluated, with the final detector trained on all available public labeled data. The pseudo-mask pipeline runs offline for the training and validation splits and derives artifact regions only for fake images, with real-image targets set to zero. The method suits settings that require both an authenticity decision and visible forensic explanations, such as auditing the basis of decisions in content moderation and forensic analysis; because explanation generation is routed by the detector prediction, its output is tied to that decision.
The authors note in the discussion that the system is computationally heavy, combining several DINOv3 backbones, Mesorch features, and multiple Qwen3-VL models, which makes it expensive and less attractive for practical deployment, and that the VLM is used mostly as an attachment to the detector: it does not participate in the authenticity decision, compare alternative hypotheses, or verify detector evidence. They also note that although GRPO improves the official simple-explanation metrics, especially normalized SLE, this does not necessarily mean the explanations become substantially more useful or accessible to ordinary users. In addition, the explanation evaluation counts only generations with the correct final answer to isolate explanation quality from detector errors, so those metrics reflect explanation quality given a correct decision, and the pseudo-masks are explicitly described as not pixel-accurate, with coverage bounded by agreement between textual evidence and open-vocabulary grounding.
