Researchers invert multi-vector visual document indices with a conditional flow-matching inverter, recovering 47% of words and 45% of sensitive tokens and ranking the source page first 98.4% of the time
Synopsis
The work frames inversion of a multi-vector visual document index as conditional document image generation, identifying the encoder, inferring page shape and restoring shuffled order from the stored vectors alone under zero side information, and redrawing the page with a conditional flow-matching inverter; on ViDoRe v3, pages inverted from the raw index recover 47% of words and 45% of sensitive tokens and rank their source page first 98.4% of the time.
Interpretation
A document page can be reconstructed from its stored multi-vector index alone: pages inverted from the raw index recover 47.4% of reference words and 45.0% of sensitive tokens at a character-level NED of 0.292, and rank their source page first among 19,252 indexed pages 98.4% of the time. Prior embedding-inversion work targets text embeddings and natural-image features, not the multi-vector page index of a visual late-interaction retriever; this work casts inversion as conditional document image generation and performs it under zero side information. Evaluated on a fixed 2,000-page set from ViDoRe v3, with inverters trained on 682,818 public pages from different sources; controls include a null model (word recall 0.001), a kNN baseline (0.234) and a VAE reconstruction upper bound (0.956), and an independent judge encoder gives 97.7% top-1.
The three facts the stored index omits are all recoverable from the index itself: the encoder is identified in all 26,000 single-page decisions over 13 public retrievers, raw-index shape is recovered for 99.2% of pages and shuffled-index shape for 99.95%. A stored index does not record which encoder produced it, the page shape, or the permutation of a shuffled index; the work supplies these through a structure filter plus reference cloud, periodicity and a set classifier, and an aspect head on the position model, so the attack needs no side information. Encoder identification is evaluated on stores built from 13 public encoders, with the reference cloud making no error in 26,000 single-page decisions; shape inference is reported on the fixed evaluation set with a true-shape control showing little change.
Two cheap protections behave differently: pooling cuts word recall to about 8% and is not inverted by the attacks tried, while shuffling is partly defeated by a position model that raises source-page top-1 from 3.8% to 93.5%. Protections had not been evaluated against inversion of visual late-interaction indices; the work shows shuffling leaves retrieval unchanged yet is weakened by order recovery, while inverting a pooled index remains open. Pooling at factors 3 and 9 gives word recall 7.6% and 8.1%, below the kNN baseline; the position model places shuffled vectors with mean errors of 1.8 rows and 4.5 columns, and after re-ordering word recall is 23.1%, sensitive-token recall 21.2% and independent-judge top-1 is 87.7%.
The attack transfers to a second retriever but content leakage does not: pages inverted from the raw index of ColQwen3.5-4.5B rank their source page first 70.2% of the time, while word recall of 7.9% stays below the kNN baseline's 22.3%. The same training and inference recipe is applied unchanged to a second public retriever fine-tuned by a different group from a different backbone, testing whether the finding is specific to one model family. The second encoder's models are trained on 346,770 pages without VDR-MEGA-2; a control inverter for the first encoder trained on the same three collections reaches 55.5% word recall and 98.9% top-1, indicating the gap lies in the encoder rather than the data.
Perspective
The result applies where the index is held apart from the pages, such as a hosted vector database, separate access tiers for the index and the documents, or index snapshots and backups; in these settings the stored index may be the only copy of the page an intruder obtains. The authors recommend that a raw index be protected like the documents it encodes and note that a shuffle is not a reliable protection. For defenders, pooling cuts word recall to about 8% under the attacks tried and is the one storage configuration not inverted here; for follow-up work, the paper provides the fixed evaluation set, indices for both encoders and the protocol, with kNN and null rows as lower bounds and the ordered index as the target to approach.
Inverting a pooled index remains open; the authors state that their failure is a failure of the methods tried, not evidence that pooling is safe, and list soft positions, differentiable assignment and greater capacity as starting points. Content metrics use the source page's OCR transcript as reference, sensitive tokens are defined by a regular-expression heuristic without a human study, and which parts of an inverted page are correct cannot be verified item by item; the authors suggest drawing several samples per index and keeping only what they share to separate high-confidence leakage from invented content. Why the second encoder leaks less cannot be attributed to input resolution, vector count or training data. In addition, this parse does not include figures, so the content of Figures 3, 4 and 19 can only be understood from the body text.
