Skip to main content
Back to timeline
arXivSource publication:

XoG recovers missing reasoning paths via graph embeddings, consistently beating untrained KGQA baselines on incomplete knowledge graphs while cutting LLM token use by up to 33%

Synopsis

The work introduces XoG (eXplore-over-Graph), a framework for multi-hop question answering over incomplete knowledge graphs that uses type-level entity-relation statistics to identify candidate relations and KG embeddings to retrieve plausible missing entities, with the LLM acting as semantic selector and reasoner inside an iterative planning-exploration-reasoning process; on WebQSP, CWQ, and the Wikidata-based BRINK benchmark, XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under incompleteness, with gains persisting across multiple LLM backbones and LLM token consumption reduced by up to 33% versus a closely related planning-based approach.

Source-provided article image: Explore-over-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness
Figure 1 ·

Figure 1: Comparison of LLM+KG reasoning paradigms.

arXiv

Interpretation

XoG recovers missing reasoning paths from incomplete knowledge graphs instead of asking the LLM to generate missing knowledge. Most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts; alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. XoG instead recovers paths from learned graph structure. The abstract states this motivation and positioning explicitly; it is a framework-level methodological claim without specific ablation numbers.

XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. It integrates symbolic statistics, embedding retrieval, and LLM semantic judgment into an iterative planning-exploration-reasoning process, rather than relying solely on graph traversal or LLM parametric knowledge. The abstract describes each component's role and how they are integrated, which is method-description-level evidence.

On WebQSP, CWQ, and the Wikidata-based BRINK benchmark, XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. It places evaluation emphasis on the incompleteness condition and reports consistent advantages over comparable methods. The abstract reports experimental results on three benchmarks with the comparison condition (no task-specific KGQA training) stated, but gives no specific numbers or sample sizes.

Gains persist across multiple LLM backbones, and XoG reduces LLM token consumption by up to 33% compared with a closely related planning-based approach. This indicates stronger LLMs alone do not resolve missing graph evidence, while also providing a quantified efficiency-side benefit. The abstract reports cross-backbone consistency and a token reduction of up to 33%, which is an abstract-level quantitative statement.

Perspective

The work targets multi-hop question answering over incomplete knowledge graphs, suited to settings where reasoning rests on structured graph evidence and reasoning paths may be broken by missing facts; its design goal is to let the LLM avoid using parametric knowledge to fill missing facts, instead recovering paths from learned graph structure. The abstract indicates the framework remains competitive on complete KGs, so it is not limited to the incomplete setting. For readers deploying QA systems over incomplete KGs or needing to control LLM token cost, XoG offers a reference planning-exploration-reasoning process and component division of labor.

The abstract does not give specific metric values per benchmark, sample sizes, or how incompleteness is constructed, nor does it state the relative contribution of type-level statistics versus embedding retrieval; cross-backbone consistency is reported but the specific backbones are not listed. Readers needing to judge applicability to their own graphs and domains will still need the experimental setup and ablation results in the full text.

Sources