Skip to main content
Back to timeline
arXivSource publication:

Evi-VN injects multimodal evidence into GNN-common hard regions via virtual class nodes, raising mean accuracy by 6.77 points across 13 MGTAB backbones

Synopsis

Evi-VN is a plug-and-play enhancer that identifies GNN-common hard regions from multi-backbone out-of-fold predictions, trains a router to flag likely hard nodes from feature-isolated multimodal evidence chains, and uses a pair of virtual class nodes to inject a calibrated class prior through source-only edges into routed nodes, leaving the backbone input schema unchanged. On MGTAB it improves mean Accuracy/F1/AUC over backbone-only by 6.77/10.58/5.61 percentage points and exceeds logit fusion by 1.10/1.27/0.74 points.

Source-provided article image: Evi-VN: Hard Region Guided Virtual Node Evidence Injection for GNN-Based Fraud Detection
Figure 1 ·

Figure 1: Overview of Evi-VN . (a) Discovery-GNN OOF errors define consensus-hard targets; the evidence router gates prior generation and calibration. (b) Original and augmented versions of one target backbone run in parallel. A calibrated prior selects and weights a source-only virtual edge, where v p v_{p} and v n v_{n} denote the positive- and negative-class virtual nodes, respectively. The finalizer combines fold-held-out graph, prior, and routing scalars. (c) Fixed class-prototype inputs become hidden virtual-node proxies under proxy alignment, bounded separation, and stop-gradient anchors ( a 0 a_{0} and a 1 a_{1} ). Discovery backbones are training-only, and evidence-chain features never enter the GNN input.

arXiv

Interpretation

The paper identifies and quantifies a GNN-common hard region: across six datasets, nodes marked hard by other backbones (wrong or near the decision boundary) show an average easy-hard accuracy gap of 0.416 and AUC gap of 0.468. Prior work targets specific graph pathologies such as camouflage, class imbalance, or homophily/heterophily; here the overlapping errors of multiple backbones are themselves used as transferable supervision, and leave-one-backbone diagnostics show the region is not one model's own errors. Out-of-fold diagnostics on six public datasets (MGTAB, TwiBot-20, Cresci-2015, YelpZip, FraudBench, TeleAntiFraud-28k); on MGTAB every target backbone shows mean accuracy/AUC gaps of 0.424/0.445 on the leave-one-out hard region, and a Random Forest predicts consensus-hard membership with 0.849 AUC.

Evi-VN makes evidence injection structural rather than feature concatenation: the router only estimates whether the graph is insufficient, the evidence processor outputs a scalar class prior, and two virtual class nodes deliver that prior through source-only directed edges to routed hard nodes, with evidence features never entering the GNN feature matrix or the final classifier. Early fusion couples GNN inputs to modality encoders, while late fusion preserves the backbone but cannot influence message passing; Evi-VN uses virtual nodes as a conditioned channel that injects sample-specific external evidence only into a routed subset, keeping backbone sampling and aggregation unchanged. Ablation shows matched stacking without virtual-node injection still trails Evi-VN by 1.08 points accuracy and 1.10 points F1; removing the prior score drops performance further (1.61 and 2.02 points), indicating gains come partly from score combination and partly from structural injection.

On MGTAB, Evi-VN reaches mean Accuracy/F1/AUC of 0.9283/0.8707/0.9769 across 13 discovery backbones, improving over backbone-only by 6.77/10.58/5.61 points and exceeding logit fusion by 1.10/1.27/0.74 points; on leave-one-out consensus-hard nodes accuracy rises from 52.97 to 66.79 (+13.82 points). Controls include a GNN probability ensemble and weighted probability/logit fusion that use the same routed prior only after message passing, so the remaining gap supports converting evidence into graph structure. Paired bootstrap 95% confidence intervals and one-sided sign-flip tests over 13 backbones; the hard-subset evaluation excludes the target backbone from hard-region discovery.

Gains transfer to unseen backbones and multiple evidence modalities: on SGC, Graph Transformer, FiLM, ResGatedGCN, CARE-GNN, and PC-GNN, mean Accuracy/F1/AUC improve by 5.18/7.62/3.27 points; cross-dataset accuracy gains range from 3.45 points on Cresci-2015 to 35.47 points on TeleAntiFraud. These backbones do not contribute to hard-region discovery, separating transferable evidence from co-adaptation to the discovery set; six router/prior pairings also keep positive gains, supporting component replaceability. Unseen-backbone transfer experiments, six router/prior pairings, and cross-dataset results spanning structured, textual, image-text, and audio-text evidence chains.

Perspective

The framework targets enhancement of existing GNN fraud detectors: users need a graph, data-side node features, an evidence chain kept separate from graph features, and several discovery backbones to construct hard-region supervision. It applies to bot, fake-review, refund-evidence, and telecom-fraud tasks where the evidence chain stays isolated from labels, splits, GNN votes, and hard annotations. Router and prior generator can be swapped per dataset while the virtual-node interface stays fixed; inference needs only the target graph learner, the router, and the prior generator for routed nodes.

MGTAB has paired tests and a five-seed finalizer check, but full-pipeline multi-seed reruns remain expensive; evidence processors are replaceable, yet current experiments cover selected tabular, language, vision, and audio encoders rather than every model family. FraudBench retains highly separable generated text and TeleAntiFraud is synthetic and transcript-dominated, so near-saturated AUC on these datasets reflects dataset construction more than broad multimodal generalization; FraudBench uses one split and one seed, its six generator graphs share real claims, and its small AUC gain over stacking is not statistically established. The YelpZip LLM prior is weak on its own, so that result should be read as Evi-VN using weak natural-language priors rather than the LLM prior solving the task. Readers should still watch how stable the hard-region definition is under different data distributions and whether shortcut-field controls in evidence-chain construction hold in more settings.

Sources