Skip to main content
Back to timeline
arXivSource publication:

RenderRank reranks documents from compressed visual tokens rendered as images, reaching 55.96 average NDCG@10 on BEIR with 16.5–35.5% fewer input tokens

Synopsis

RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.

AI-generated editorial illustration: RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Interpretation

It introduces RenderRank, a reranker that scores query-dependent relevance from compressed visual tokens instead of text tokens, with document images encoded in advance independently of the query so that scoring feeds only the textual query and the visual representation to the language backbone. Visual text compression had mainly been used for generation tasks such as question answering and summarization, and image–text reranking work focused on precomputed visual features or listwise single-pass ranking; RenderRank carries this representation into per-candidate relevance scoring while keeping the query in text form. The paper specifies the rendering procedure (Roboto Regular 12pt, line spacing 1.0, 96 DPI, 896-pixel width, height in 32-pixel increments, continuation images for long documents) and the two training objectives (score-level MSE distillation and query-local InfoNCE), evaluated on 11 BEIR datasets and four long-document datasets.

Each training stage contributes measurable gains: cross-modal relevance distillation raises average NDCG@10 with image inputs from 50.20 to 54.86, and query-local relevance discrimination lifts it further to 55.96. The ablation shows that changing the input modality alone does not preserve a text reranker's scoring ability; the model first matches teacher scores and then separates positive from negative documents within the same query's candidate set. Table 6(a) compares no stages, stage 1 only, and both stages; training uses 1.57M records expanded into 6.28M query–document pairs for distillation and RLHN-100K organized into candidate sets for stage 2, with the vision encoder and feature mergers frozen and LoRA applied only to the text decoder.

On BEIR it reaches an average NDCG@10 of 55.96, outperforming all evaluated text-based rerankers below 4B parameters and some larger models, at 290.07 average input tokens per pair. It exceeds zerank-2-reranker and LightOn-rerank-PW-4B by 7.66% and 7.26% in average NDCG@10 respectively while using fewer tokens than the evaluated text-based rerankers. Table 2 reports per-dataset scores and average token counts across 11 BEIR datasets against baselines from 150M to 4.5B parameters; each BEIR query reranks the top 100 documents retrieved by BM25, with nDCG@10 as the metric.

On long documents it reaches an average NDCG@10 of 88.27 with roughly half the average input token count of the evaluated text-based rerankers, and the highest throughput among the compared models on all four datasets. On MLDR it scores 99.74 NDCG@10 at about 4.20K average input tokens, and on the three LongEmbed datasets it uses 53–57% fewer input tokens than the most token-efficient text baseline, averaging 4.51 PPS versus 2.66 PPS for the fastest baseline. Table 4 reports NDCG@10, token counts, and PPS for MLDR, 2WikiMQA, QMSum, and SummScreenFD; fixed-length-budget experiments at 2K, 4K, and 8K show RenderRank highest at all three, with 97.9 at 2K already exceeding the 95.8 and 97.3 that gte-reranker and LightOn-PW-4B reach at 4K.

Perspective

The work targets reranking text candidate documents against textual queries in settings where candidate documents can be rendered and their visual embeddings cached in advance; the paper explicitly assumes document image visual embeddings have been computed beforehand, reporting 76.83 PPS with caching versus 37.31 PPS when visual encoding is included in the measurement. Evaluation covers 11 BEIR datasets plus the long-document datasets MLDR, 2WikiMQA, QMSum, and SummScreenFD, with queries kept as text and documents rendered as images. Rendering configuration (font size, line spacing, image dimensions) sets the compression level; the paper selects 12pt with line spacing 1.0 based on question answering and summarization performance and token efficiency, and notes the trade-off is adjustable: moving from 10pt to 14pt raises average input length from about 236 to 370 tokens and lowers throughput from about 94 to 63 PPS while NDCG@10 rises from 55.04 to 56.43. For retrieval system builders aiming to cut reranking compute or fit longer documents into a fixed length budget, these results offer a reproducible starting point.

The effect of rendering configuration varies across models and tasks: the paper reports that at 12pt, increasing line spacing from 1.0 to 1.2 improves Qwen's Multi-News score but reduces token savings from 42.66% to 27.92%, while Gemma at 12pt with line spacing 1.2 uses more document tokens than the text baseline without reaching its performance. Estimated TFLOPs and measured throughput do not always align: RenderRank's estimated TFLOPs per pair exceed LAMAR-600m's, yet its throughput of 76.83 PPS exceeds the 67.55 PPS measured for LAMAR-600m, and computational cost differences also reflect backbone size and attention architecture. Results come from a single training run per configuration with the random seed set to 42. How caching and I/O costs of precomputed visual embeddings factor into real serving, and how compressed visual representations perform on corpora and languages beyond the BEIR and long-document datasets, remain open questions worth watching.

Sources