Skip to main content
Back to timeline
arXivSource publication:

G2MLP stratifies distillation supervision by Ollivier–Ricci curvature, achieving the best graph-free student accuracy on all six node-classification benchmarks and a 4.02-point gain on Citeseer

Synopsis

The authors propose G2MLP, a training-time GNN-to-MLP distillation framework that uses Ollivier–Ricci curvature to allocate supervision between prediction-level and representation-level alignment, correcting spectral underfit on sparse graphs and spectral overfit on dense graphs; the graph-free MLP student achieves the best accuracy on all six node-classification benchmarks and transfers without architectural change to Graph Transformer teachers and link prediction.

Source-provided article image: Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs
Figure 1 ·

Figure 1: Two spectral failure modes of GNN-to-MLP distillation. (a) Singular-value spectra: GLNN exhibits spectral underfit on sparse graphs and spectral overfit on dense graphs. (b) Rank-direction sketch. (c) Per-dataset rank gap.

arXiv

Interpretation

The paper characterizes two spectral failure modes of GNN-to-MLP distillation: on sparse graphs the student has lower effective rank than the teacher and misses high-energy teacher directions near boundary regions, termed spectral underfit; on dense graphs the teacher collapses to a low-rank geometry through aggregation while the feature-only student retains directions the teacher weakly supports, termed spectral overfit. Prior methods mainly transfer node-wise predictions or reweight by teacher confidence without specifying where the student should preserve the teacher's graph-induced geometry; this work locates the mismatch in the spectrum of the representation space and ties it to boundary versus interior graph regions. Grounded in an analysis of the GNN kernel versus the MLP kernel: Lemma 3.1 states that the graph-dependent aggregation component is generated through paths in -hop receptive fields, and Proposition 3.2 gives an upper-bound relation between ORC and local random-walk neighborhood alignment difficulty under a Lipschitz assumption; the authors state these are structural upper bounds and do not claim curvature exactly estimates the kernel spectrum.

The paper derives an energy-weighted alignment objective showing why pointwise KL distillation is insufficient: two students can match the same logits while inducing different neighborhood representation geometries, so an additional objective sensitive to neighborhood-level teacher-student alignment is required. It extends the distillation target from node-wise divergences to energy-weighted alignment in the teacher kernel's eigenbasis, arguing that high-energy teacher directions should receive larger weight. The target is impractical to use directly because it requires storing and diagonalizing the full teacher kernel; the paper treats it as a principled target and approximates it with local, curvature-guided alignment.

G2MLP uses ORC as a local proxy to allocate supervision: low-curvature boundary nodes receive stronger prediction-level supervision to correct spectral underfit, while high-curvature interior nodes receive stronger representation-level neighborhood alignment to reduce spectral overfit; the representation term uses a closed-form diagonal Gaussian surrogate, and all graph-dependent quantities are used only during training. It ties reweighting to a graph-topology quantity (curvature) rather than teacher confidence, and applies sign-opposite monotone weights to the logit and feature components simultaneously. Ablation shows alignment under uniform weights is neutral relative to GLNN (Cora 80.86 vs. 80.26, Citeseer 71.07 vs. 71.22, Pubmed 75.22 vs. 75.59); single-side curvature stratification yields most of the improvement (roughly 1.5–2.2 points), and both sides together are best (Cora 82.54, Citeseer 74.51, Pubmed 78.17), consistent with the sign-opposite stratification predicted in Section 3.2.

Across six node-classification benchmarks, G2MLP achieves the best graph-free accuracy in the transductive setting, improving over the strongest baseline by 0.16–1.66 points, and surpasses the GraphSAGE teacher on five of six datasets (up to +4.02 points on Citeseer); in the inductive setting it improves by 2.07–2.69 points on Cora, Citeseer, Pubmed, and A-computer. The inductive gains are substantially larger than the transductive ones, suggesting curvature-stratified alignment transfers better to unseen nodes than node-wise distillation; the same recipe transfers to Graph Transformer teachers and link prediction without architectural change. Reports mean and standard deviation over 10 random seeds; on ogbn-Arxiv the student still falls short of the teacher by 6.16 points, which the authors attribute to the known difficulty of feature-only inference on that large-scale graph.

Perspective

The results target homophilic node-classification benchmarks (Cora, Citeseer, Pubmed, A-computer, A-photo, ogbn-arxiv), where neighborhood Gaussians are statistically reliable and the diagonal moment surrogate is well-justified; the deployed form is a standard MLP requiring no neighborhood fetching, sampling, or graph traversal at inference, making it suitable for latency-sensitive settings with limited graph access. Curvature is computed once per graph as preprocessing and cached, so the overhead amortizes across all training runs and deployment; for graphs where exact ORC is prohibitive, Sinkhorn-regularized optimal transport or Forman–Ricci curvature can substitute. The method also applies to Graph Transformer teachers and to link prediction, the latter requiring per-dataset validation of the sign of the feature-side weight.

The theoretical results (Lemma 3.1 and Proposition 3.2) are upper bounds and structural statements; no matching lower bounds are established, and it is not proven that curvature-stratified regimes are necessary for closing the spectral mismatch. Which graphs and which teacher kernels genuinely require curvature-adaptive treatment remains an open question. The diagonal Gaussian neighborhood surrogate is exact only when the underlying neighborhood distribution is itself Gaussian with diagonal covariance, and may under-represent the alignment cost for multimodal or strongly heterophilic neighborhoods; mixture-of-Gaussian or low-rank covariance extensions remain untested. Empirical evaluation is restricted to homophilic benchmarks, so behavior on strongly heterophilic graphs (e.g., Texas, Wisconsin, Chameleon, Squirrel) and whether sign-flipped weighting is needed await systematic study. In addition, the student still falls short of the teacher by 6.16 points on ogbn-arxiv, leaving the feature-only gap on large-scale graphs unclosed.

Sources