CorpusMap links a corpus into an entity graph, raising answer quality by 6.4–11.7 points and cutting input tokens by 34–57% across 7 models and 3 benchmarks
Synopsis
The work introduces CorpusMap, an offline-built, entity-anchored navigation layer that renders each recurring cross-document entity as a source-attributed Entity Page linked to every document mentioning it, and shows across 7 models and 3 multi-document QA benchmarks (EnterpriseRAG-Bench, WixQA, HERB) that it improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers.
Interpretation
CorpusMap organizes a corpus as a bipartite graph between entities and documents: each retained cross-document entity becomes an Entity Page that consolidates what documents state about it, attributes each fact to its source, and links to every document that refers to it. Earlier RAG and graph-based approaches such as GraphRAG and HippoRAG keep entity relations inside the retriever, so the generator sees only a fixed context; agent-facing structures such as LLM Wiki and Corpus2Skill either let the model decide what becomes a page or assign documents to a single topical branch. CorpusMap instead exposes entity–document links directly to the agent, and those links are determined by the corpus itself rather than by any query. The paper gives a formal definition (the bipartite graph of Equation 1), a four-stage offline construction protocol (Cataloging, Extraction, Resolution, Rendering), and states that Entity Pages are stored as files alongside raw documents so the agent reads and searches them with the same shell commands.
Across 7 models and 3 benchmarks, CorpusMap beats raw-corpus agentic search on both answer quality and retrieval quality while consuming fewer tokens on average. The abstract and main text report a 6.4 to 11.7 point gain in overall quality over raw-corpus agentic search, with 34% to 57% fewer input tokens on average for GPT models; a paired bootstrap test shows the gains are significant against every baseline for every GPT model, with all 95% confidence intervals above zero. The main table covers four GPT models (GPT-5.5, Luna, Terra, Sol) and the trend is reproduced with DeepSeek-V4-Pro and MAI-Thinking-1; benchmarks are 80 multi-document questions from EnterpriseRAG-Bench, 79 multi-document questions from WixQA, and 238 content-based questions from HERB, over corpora of 2,819, 6,221, and 6,365 documents respectively.
CorpusMap outperforms 4 alternative navigation layers and also beats retrieval-based methods including BM25, dense retrieval, HippoRAG, and GraphRAG. The paper notes that the other organization schemes (Document Page, Group Page, LLM Wiki, Corpus2Skill) do not consistently improve over raw corpus, indicating that merely adding a navigation layer does not guarantee gains, while retrieval-based methods are limited by what a single retrieval step returns. The main and generalization tables report document recall, correctness, completeness, factuality, content, and tokens for each method across the three benchmarks; the retrieval comparison is run with GPT-5.5 and Luna.
The map can be built offline, reused across models, updated incrementally, and constructed with off-the-shelf tools instead of an LLM. The paper reports that a map built by the least expensive LLM (costing $74.65) improves overall quality over raw corpus for every answering LLM; that incrementally extending the map in chronological order saves a substantial fraction of full-rebuild tokens while the final map performs comparably to a full rebuild; and that a map built with off-the-shelf components such as GLinker and GLiNER models is comparably effective to the LLM-built map at lower cost. The reuse experiment spans four map builders (GPT-5.5, Sol, Luna, DeepSeek) and four answering models; the incremental experiment orders the EnterpriseRAG-Bench corpus by its provided timestamps; the off-the-shelf experiment also improves over raw corpus with the open-weight Qwen3.8-27B.
Perspective
The result targets large heterogeneous corpora where evidence is spread across multiple documents and documents are generated around recurring entities such as people, projects, or products, for example enterprise issue trackers, shared drives, and chat channels; the paper validates it on three benchmarks (EnterpriseRAG-Bench, WixQA, HERB) with fixed corpora, and states that the map can be built offline, reused across models, updated incrementally as the corpus grows, and constructed with off-the-shelf components such as GLinker instead of an LLM. For questions grounded in a single document, the paper reports CorpusMap still outperforms raw corpus with fewer tokens.
The entity–entity edge analysis shows such edges only shorten paths the bipartite map already provides: the agent follows them in only 11–19% of questions, 71% of those moves lead to no new gold document, and they enlarge Entity Pages by 27–52% on average and raise input tokens by up to 28%, so which corpora need richer entity relations remains an open question. The candidate file path experiment shows that without candidates the agent still beats raw corpus but spends far more tokens exploring, and how candidate count and quality behave across different corpora is not yet fully characterized. The ethics statement also notes that consolidating information about one entity into a single page may aggregate otherwise dispersed sensitive information or expose documents to users without access, so real deployment would need to build and serve the map according to corpus access permissions. In addition, this reading covers the full paper text and the homepage abstract, but not the raw figure data or every prompt detail in the appendix.
