Public articles linked to the same research event.
arXiv The authors introduce the Graph Transformer Language Model (GTLM), which lets a pretrained LLM process graph topology natively by injecting graph-aware attention biases directly into its attention modules, adding only about 0.015% structure-related parameters; they prove the bidirectional attention prefix is permutation-equivariant over nodes and that GTLM reduces exactly to the pretrained model when no graph is present, and report that it matches or exceeds domain-specific state-of-the-art models on text-attributed graph benchmarks, GraphRAG on WebQSP, and molecular benchmarks, meaningfully improves over strong baselines on GraphQA, and keeps needle-in-a-graph accuracy flat from 1k to 64k tokens and 4x past the training length while an identically trained flat-text baseline collapses.
The authors introduce the Graph Transformer Language Model (GTLM), which lets a pretrained LLM process graph topology natively by injecting graph-aware attention biases directly into its attention modules, adding only about 0.015% structure-related parameters; they prove the bidirectional attention prefix is permutation-equivariant over nodes and that GTLM reduces exactly to the pretrained model when no graph is present, and report that it matches or exceeds domain-specific state-of-the-art models on text-attributed graph benchmarks, GraphRAG on WebQSP, and molecular benchmarks, meaningfully improves over strong baselines on GraphQA, and keeps needle-in-a-graph accuracy flat from 1k to 64k tokens and 4x past the training length while an identically trained flat-text baseline collapses.
The authors introduce the Graph Transformer Language Model (GTLM), which lets a pretrained LLM process graph topology natively by injecting graph-aware attention biases directly into its attention modules, adding only about 0.015% structure-related parameters; they prove the bidirectional attention prefix is permutation-equivariant over nodes and that GTLM reduces exactly to the pretrained model when no graph is present, and report that it matches or exceeds domain-specific state-of-the-art models on text-attributed graph benchmarks, GraphRAG on WebQSP, and molecular benchmarks, meaningfully improves over strong baselines on GraphQA, and keeps needle-in-a-graph accuracy flat from 1k to 64k tokens and 4x past the training length while an identically trained flat-text baseline collapses.
The authors introduce the Graph Transformer Language Model (GTLM), which lets a pretrained LLM process graph topology natively by injecting graph-aware attention biases directly into its attention modules, adding only about 0.015% structure-related parameters; they prove the bidirectional attention prefix is permutation-equivariant over nodes and that GTLM reduces exactly to the pretrained model when no graph is present, and report that it matches or exceeds domain-specific state-of-the-art models on text-attributed graph benchmarks, GraphRAG on WebQSP, and molecular benchmarks, meaningfully improves over strong baselines on GraphQA, and keeps needle-in-a-graph accuracy flat from 1k to 64k tokens and 4x past the training length while an identically trained flat-text baseline collapses.