KGPFN uses in-context learning to give knowledge graph foundation models the best average MRR across 57 graphs
Related research and updatesSynopsis
The authors propose KGPFN, a knowledge graph foundation model built on a Prior-Data Fitted Network that learns relation representations via message passing on relation graphs, extracts multi-scale local context from the intermediate head representations of a multi-layer NBFNet, and builds relation-specific global context from positive and negative examples of the query relation, aggregating this context with feature-level and sample-level attention; after multi-graph pretraining it combines structural representations with labeled contextual evidence without inference-time parameter updates, achieving the best average MRR on 57 knowledge graphs both without and with fine-tuning, with context sensitivity analyses highlighting the value of negative context examples.
Figure 1: (a) Local context. The same structural cue may support different conclusions depending on the query-specific neighborhood. (b) Global context. Global context summarizes how a relation is typically instantiated across many examples and their neighborhoods.
arXivInterpretation
KGPFN brings inference-time in-context learning to knowledge graph foundation models, so prediction conditions on both the local neighborhood of the query entities and a global context that summarizes how the query relation behaves across many instances. Most existing methods focus on relation-level universality, leaving in-context learning, the other pillar of foundation models, largely unexplored for KG reasoning; KGPFN explicitly incorporates structured, heterogeneous graph context into conditioning. The abstract states that existing methods leave in-context learning 'largely unexplored for KG reasoning' and describes KGPFN as conditioning on both local neighborhood and global context.
The model obtains transferable relation representations through message passing on relation graphs and extracts multi-scale local context from the intermediate head representations of a multi-layer NBFNet. It couples relation-level representation learning with multi-scale local structural context rather than relying on a single level of relation representation. The abstract describes that KGPFN 'learns relation representations by message passing on relation graphs and extracts multi-scale local context from the intermediate head representations of a multi-layer NBFNet'.
KGPFN builds relation-specific global context from positive and negative examples of the query relation together with their local structural representations, and aggregates this context with feature-level and sample-level attention. It treats labeled positive and negative contextual evidence as a source of global context and aggregates it through two levels of attention; the abstract's context sensitivity analyses further 'highlight the value of negative context examples'. The abstract states that global context is built from positive and negative examples and their local structural representations and aggregated with feature-level and sample-level attention; context sensitivity analyses support the value of negative context examples.
Through multi-graph pretraining, KGPFN learns to combine structural representations with labeled contextual evidence without inference-time parameter updates, and achieves the best average MRR on 57 knowledge graphs both without and with fine-tuning. It replaces inference-time parameter updates with PFN-style feed-forward in-context learning and covers both non-fine-tuned and fine-tuned settings in a large multi-graph evaluation. The abstract reports that 'On 57 knowledge graphs, KGPFN achieves the best average MRR both without and with fine-tuning' and that it operates 'without inference-time parameter updates'.
Perspective
The work targets researchers and practitioners who need knowledge graph reasoning over unseen entities and relations, in settings centered on link-prediction-style queries where positive and negative examples of the query relation can be supplied as context; its multi-graph pretraining setup means the method is intended to operate when multi-graph training resources are available, and it can be used without inference-time parameter updates.
The abstract does not give per-dataset results, the specific scale of context construction, or ablation details, so the magnitude of the value of negative context examples and differences across graph types still need confirmation in the full text; the composition and scale of the graph collection used for multi-graph pretraining are also not expanded in the abstract, which is a direction readers can continue to watch.
