Skip to main content
Back to timeline
arXivSource publication:

ColNanoVDR distills multi-vector visual document retrieval queries into a 149M text-only student, keeping about 95% of NDCG@5 and encoding queries 26x faster on one CPU thread

Synopsis

The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.

AI-generated editorial illustration: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport

Interpretation

It is the first framework to bring document-free query-side distillation to multi-vector visual document retrieval, aligning only the teacher's query tokens and encoding or reading no page during training. NanoVDR had done document-free distillation only for single-vector retrievers, while the standard recipe for multi-vector students is score distillation, which requires caching every training page (the paper estimates about 2 TiB of cached page tokens for one million ColVec1.1 training pairs). The paper gives the method derivation and experiments across five teachers on ViDoRe v1-v3, and reports that OTW reads 12.6x less cached teacher data than score distillation (342.7 versus 27.3 GiB).

It proposes the OTW objective, which aligns the student's and teacher's query token sets by entropic optimal transport and predicts a weight for each student token from the query with a linear head, requiring no correspondence between the two tokenizations. Prior optimal-transport distillation aligned label distributions, output distributions, or in-batch features, whereas here what is transported are the two token-embedding sets from which the retrieval score itself is computed; the weights also differ from prior work that sets them from corpus statistics or relevance labels, being learned as the student-side marginal of the transport plan. The paper provides a theorem and proof (Theorem 1 and Appendix A) and compares against InfoNCE, Listwise KL, Coverage, and OT-uniform under identical training; on ColQwen3.5, OTW matches the strongest document-dependent baseline Listwise KL on v1 and v3 and trails on v2 (60.0 versus 61.2), and on Tomoro-ColQwen3-8B it leads on all three benchmarks, by 5.0 points on v3.

It proves that the alignment cost bounds the MaxSim score difference on every page, so aligning queries alone suffices in principle to control retrieval scores. MaxSim takes a maximum over page tokens, so it was not obvious that query-side alignment controls the score; the theorem turns the quantity OTW minimizes into a sufficient condition on the score discrepancy for pages never seen in training. Appendix A gives the full proof, and Appendix E measures the bound on 4,735 evaluation queries, finding it holds and is not vacuous but is loose, exceeding the worst-case discrepancy it certifies by about a factor of eight; the paper explicitly calls it sufficient rather than necessary.

Across five teachers it demonstrates fidelity and efficiency: the 149M text-only students retain about 99% of their teacher's NDCG@5 on v1 and 93.0%-96.5% on the harder v2 and out-of-domain v3 enterprise collections, while encoding a query in 87 ms on one CPU thread, 26x faster than the ColQwen3.5 teacher. Only the query encoder is replaced while the teacher's document encoder and page index stay unchanged, so no collection is re-indexed and the method is complementary to index-side compression. Tables 1 and 2 and Figure 1 give per-teacher numbers; a capacity ablation shows retention rising monotonically from Ettin-32M to Ettin-400M; and in the index-compression experiment retention falls only from 95.8% to 95.3% at pool factor 9.

Perspective

The result targets settings that already deploy multi-vector visual document retrieval and want lower query-encoding latency without re-indexing: the teacher's document encoder and page index are kept as they are, and the student plugs into an existing late-interaction engine with weighted MaxSim, so it composes with index-side token pooling. It applies to ViDoRe-style page-image retrieval with predominantly English query workloads, and to training pipelines that change teachers often and want to avoid re-caching pages. The capacity ablation gives retention curves from 32M to 400M, which helps choose a student size against a latency budget.

The paper states that all students come from a single Ettin encoder family and that every configuration is a single run with a fixed seed and no reported variance; on ColQwen3.5 the margins between document-free objectives are at most two points, and those on v1 lie within the range a single run cannot resolve. The objective comparison covers two teachers only, and the remaining three teachers were trained with OTW alone, so whether score distillation and the fixed-weight alternatives order the same way under them is untested. Training queries read no pages but all come from a paired set, and training on unpaired query logs is not exercised. Only the query encoder is compressed: index size and storage remain the teacher's, and search cost falls only through the shorter query. The theoretical guarantee is a sufficient condition, measured to be loose by about a factor of eight, and it does not distinguish between weightings of the student measure. Evaluation is predominantly English and results are not broken down by language.

Sources