Co-LMLM interleaves vector retrieval queries with next-token prediction during pre-training, reaching lower perplexity at 360M than models trained on 40x more data and SimpleQA performance in line with gpt-4o-mini and above Claude Sonnet 4.5
Synopsis
The work proposes continuous-query limited memory language models (Co-LMLM), an LLM that interleaves flexible vector retrieval queries with next-token predictions and is pre-trained to copy knowledge returned from the KB rather than memorize it, using a scalable approach that jointly trains a knowledge-externalizing LLM, induces its knowledge base, and learns an expressive continuous retrieval mechanism; across pre-training at multiple model scales, Co-LMLM outperforms prior knowledge-externalizing and vanilla LLMs in both perplexity and factual precision, and at 360M scale it achieves lower perplexity than models pre-trained on 40x more data and SimpleQA-verified performance in line with gpt-4o-mini and higher than Claude Sonnet 4.5.
Figure 1: Knowledge separation across three regimes. A standard LLM with RAG retrieves over external documents but keeps factual knowledge in its parameters ( left ); Rel-LMLM externalizes facts to a relational KB queried with an explicit decoded query ( middle ); Co-LMLM externalizes facts to an unstructured index, retrieved directly from the model’s hidden state ( right ). Bottom: some of the advantages and limitations of each regime.
arXivInterpretation
Introduces Co-LMLM, a language model that interleaves flexible vector retrieval queries with next-token predictions and is pre-trained to copy knowledge returned from the KB rather than memorize it in parameters. Relative to prior knowledge-externalizing and vanilla LLMs, the design shifts knowledge use from parametric memorization to retrieval-then-copy, which the authors position as a path to higher performance at smaller scales, controlled knowledge use, and greater model transparency. The abstract states the design claim and reports cross-scale pre-training comparisons, but does not give dataset composition, the form of the retrieval queries, or how transparency is measured.
Proposes a scalable pre-training approach that jointly trains a knowledge-externalizing LLM, induces its knowledge base, and learns an expressive continuous retrieval mechanism. Knowledge-base induction and continuous retrieval learning are folded into the same pre-training procedure rather than training a retriever separately against a fixed knowledge base. The abstract describes the procedure as jointly trained but provides no training-stage breakdown, knowledge-base size, or ablation details.
Across pre-training at multiple model scales, Co-LMLM outperforms prior knowledge-externalizing and vanilla LLMs on both perplexity and factual precision. Compared with both the existing knowledge-externalizing line and vanilla pre-training, the advantage appears on two metrics at once rather than a single metric. The abstract reports this as a consistent result across multiple model scales, but lists no per-scale numbers, comparison model roster, or evaluation protocol.
At 360M scale, Co-LMLM achieves lower perplexity than models pre-trained on 40x more data, and SimpleQA-verified performance in line with gpt-4o-mini and higher than Claude Sonnet 4.5. Concretizes the payoff of the knowledge-externalizing route as a perplexity advantage for a small model over models trained on far more data, plus factual precision on SimpleQA comparable to a larger proprietary model. The abstract gives the specific 360M scale, the 40x data comparison, and the SimpleQA model comparison, but no specific scores, confidence intervals, or evaluation sample sizes.
Perspective
The result addresses the knowledge-externalization route at pre-training time: it applies to research and engineering settings that want lower perplexity and higher factual precision at smaller model scale, with knowledge use that a retrieval mechanism can control and expose. The abstract's validation spans multiple model scales, with the 360M scale providing the perplexity comparison against models trained on 40x more data and the SimpleQA comparison against gpt-4o-mini and Claude Sonnet 4.5, so the scope of the conclusion is bounded by those scales and evaluation settings. For a reader, this means Co-LMLM can be treated as an architectural direction that introduces continuous retrieval queries and knowledge-base copying into pre-training, as an alternative or complement to purely parametric memory.
The abstract does not state the knowledge base's size, construction, or update procedure, nor the concrete form of the continuous retrieval queries or their training signal. Specific perplexity and factual-precision numbers, the comparison model roster, and the SimpleQA evaluation samples and judging method are not given in the text; the specific values of the multiple model scales are not listed apart from 360M. The abstract also does not discuss inference-time retrieval latency or knowledge-base maintenance cost, nor how the method performs in domains or long-tail facts outside the knowledge base's coverage. These are areas the abstract leaves unexpanded, so readers making architecture decisions should return to the full text to check experimental settings and ablation results.
