MAGIC combines a topological prior with drift-aware distillation for replay-free graph few-shot class-incremental learning, raising 5-shot mean accuracy by 5.48 points
Related research and updatesSynopsis
The authors propose MAGIC, a replay-free graph few-shot class-incremental learning framework that pairs a frozen graph representation backbone with closed-form analytic continual learning, learns a topological prior from the base graph characterizing both homophilous and heterophilous relations and injects it through Potts Markov random field inference to ease novel-class overfitting, and transfers previous predictions of affected historical nodes to their updated representations via drift-aware analytic distillation; across five datasets and eight baselines, the 5-shot setting improves Mean Accuracy and Final Accuracy by 5.48 and 9.33 percentage points on average and reduces Performance Drop by 10.78 points, with substantially less training time.
Figure 1: On the base stage, MAGIC learns a topological prior. In novel sessions, limited labeled supports first update an initial classifier, while drift-aware distillation protects historical nodes affected by cross-session edges. Finally, MAGIC extracts a subgraph around novel supports and performs Potts MRF inference with the topological prior to adjust the pseudo-labels predicted by the initial classifier, and subsequently updates the classifier using these adjusted labels.
arXivInterpretation
MAGIC combines a frozen graph representation backbone with analytic continual learning, replacing gradient-based classifier updates with closed-form updates over accumulated sufficient statistics, thereby mitigating catastrophic forgetting without replaying historical labeled data. Existing GFSCIL methods largely rely on pseudo-incremental meta-training, prototype memories, or pseudo-label calibration, whereas MAGIC recasts classifier learning as a closed-form analytic solution and injects hard labels, distillation targets, and MRF soft supervision through one shared set of sufficient statistics. The paper provides a unified Gaussian working-likelihood derivation showing the three supervision sources are mathematically compatible, and reports MA, FA, and PD over five datasets and eight baselines as means and standard deviations of five runs.
MAGIC estimates a class-balanced relational prior from the base graph that links the cosine similarity of connected node representations to their label agreement, and injects it through Potts Markov random field inference to generate additional supervision for novel classes. The prior depends only on label agreement rather than class identity, so it is invariant to class permutation and can express both homophilous and heterophilous relations, going beyond prior methods that operate mainly in node-feature space and disregard graph topological patterns. Ablation shows removing the topology prior drops Mean Accuracy from 63.50 to 53.69 and Final Accuracy from 49.81 to 37.29, the largest effect of the two components; hyperparameter analysis indicates near-insensitivity to the number of prior bins.
MAGIC introduces drift-aware analytic distillation, measuring the representation drift of historical nodes caused by cross-session edges and transferring the previous classifier's predictions on the old graph to their updated representations according to a normalized distillation mass. Theorem 1 shows that a fixed graph encoder does not imply fixed analytic statistics on a dynamic graph, indicating that freezing the encoder alone does not prevent historical representation drift, a reverse effect that prior GFSCIL work has largely not modeled. Ablation shows disabling distillation lowers Mean Accuracy to 61.88, Final Accuracy to 48.02, and raises PD to 36.16; hyperparameter analysis reveals a retention-plasticity trade-off in the distillation mass.
Across five node-classification benchmarks and eight baselines, MAGIC generally outperforms compared methods under the 1-, 3-, and 5-shot settings while substantially reducing training time. Under 5-shot, it improves MA by 5.48 points and FA by 9.33 points on average and reduces PD by 10.78 points on average relative to the best baselines; under 1- and 3-shot it attains the highest MA and FA on every dataset, with advantages growing as more supports are provided. Results cover CoraFull, Coauthor CS, Amazon Computers, ogbn-arxiv, and WikiCS, reported as means and standard deviations over five runs; on Amazon Computers MA is slightly below GFCIL while FA and PD are best.
Perspective
The work targets transductive or dynamically visible GFSCIL settings: MRF inference requires access to unlabeled nodes and edges surrounding the support set, so it fits best where the relevant neighborhood structure is available. The method is compatible with parameter-free propagation operators and frozen pretrained graph encoders, allowing backbone choice by graph characteristics, for example PPR performing best on the two citation networks, RGVT strongest on Coauthor CS, and S2GC achieving the highest accuracy on Amazon Computers and WikiCS. For practitioners who want to avoid replaying historical labeled data while continually admitting new classes, MAGIC offers a combination of closed-form updates and topological prior injection, with released code and per-dataset hyperparameters.
The paper states two scope limitations: first, it does not explicitly determine whether the topology prior learned in the base session remains beneficial in each novel session, and when novel-class relation patterns differ substantially from the base session the transferred prior may give limited or even misleading supervision, suggesting future session-level prior-validity estimators; second, MRF inference depends on unlabeled nodes and edges around the support set, making it harder to apply in strictly inductive, privacy-constrained, or topologically incomplete scenarios. Hyperparameter analysis also shows a retention-plasticity trade-off in the distillation mass, where excessive distillation suppresses novel-class learning, so that balance point merits attention in deployment.
