Skip to main content
Back to timeline
arXivSource publication:

FocusVTC renders long text as low-DPI pages and zooms only the regions reasoning needs, scoring 87.4 on RULER v1 at roughly 2.9x compression versus 57.5 for Glyph

Synopsis

FocusVTC introduces an adaptive-resolution visual text compression framework that keeps global coverage with low-DPI page images and uses tool calls to re-read reasoning-relevant regions at aligned 144 DPI, trained with multi-resolution supervised fine-tuning and GRPO on 29.4K Reasoning-Evidence Localization chain-of-thought examples that link reasoning traces to page indices and bounding boxes, reaching 87.4 on RULER v1 at 72 DPI (versus 57.5 for Glyph), 56.40 on LongBench (versus 55.86 for its text-input backbone), a 13.91-point MRCR macro-average gain, and a 51.19 VTCBench macro-average, while MMMU rises from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

AI-generated editorial illustration: FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

Interpretation

It proposes an adaptive-resolution visual text compression framework that combines compressed low-DPI global views with tool-mediated access to high-resolution content for selected regions, breaking the fixed-resolution tie between compression rate and legibility. Prior visual text compression work such as Glyph, DeepSeek-OCR, VisInContext, and VIST renders at fixed resolution, so local legibility stays tied to the global visual-token budget; FocusVTC lets global coverage and precise local reading use different resolutions. The paper formalizes the setup (low-DPI global views, an Enhance_Region(page, bbox) tool, an aligned 144 DPI enhancement source) and shows in a 48-144 DPI RULER sweep that it is far less sensitive to initial resolution than fixed-resolution baselines, reaching 87.38/75.94 on v1/v2 at 72 DPI versus 37.40/44.64 for the same backbone without crops.

It constructs 29.4K Reasoning-Evidence Localization chain-of-thought examples (REL-CoT) linking reasoning traces to page indices and normalized bounding boxes, and trains in two stages with multi-resolution REL-SFT and GRPO. REL-CoT is built by a three-stage cross-model pipeline: DeepSeek V4 Flash filters out questions answerable from common knowledge, Gemini 3.5 Flash produces reasoning traces with page indices and boxes, and GPT-5 mini independently verifies whether the selected regions support the reference answer; REL-SFT jointly supervises reasoning, evidence localization, and answer generation across seven rendering resolutions, while GRPO learns when and where to enhance and when to stop, without a separate continual-pretraining stage. The paper reports 29,411 REL-CoT examples expanding to 205,877 candidate SFT samples; ablations show REL-SFT gives a useful initialization for GRPO (LongBench 49.30 to 56.40) and that after GRPO FocusVTC surpasses the w/o GRPO+tools variant on all nine aggregates (RULER v1 from 22.98 to 87.38).

It achieves strong long-context results with compressed visual input: 87.4 on RULER v1, 56.40 on LongBench, a 13.91-point MRCR macro-average gain, and a 51.19 VTCBench macro-average, while preserving general multimodal capabilities. On LongBench, 56.40 slightly exceeds the text-input backbone Qwen3.5-9B at 55.86 and gains 20.54 points over the same backbone at fixed 72 DPI; MRCR two/four/eight-needle averages are 60.76/45.21/30.71 versus 42.63/25.97/16.57 for Glyph; VTCBench shows the best Retrieval and Memory averages. Gains concentrate on QA and synthetic retrieval where a small number of passages determine the answer, summarization changes little, and few-shot results are mixed; all six general multimodal benchmarks improve over Qwen3.5-9B (MMMU 65.12 to 66.73, MME 2424.02 to 2457.62).

It provides quantitative evidence on rendering fidelity and efficiency: font and point-size choices affect downstream scores, 72 DPI is the most balanced point between compression and performance, and online end-to-end latency drops substantially. In a sweep over 15 fonts and 8 point sizes using randomized-text character error rate, DejaVu Sans and Verdana first meet the 5% CER criterion at 9 pt, and DejaVu Sans has the lowest visual-token cost among qualifying fonts (8,368.4 versus 8,417.8 per 32K-token context) and reduces CER by 0.30 percentage points on confusable sequences; under matched settings DejaVu Sans improves LongBench from 54.93 to 56.40. At 72 DPI on RULER v1, FocusVTC uses an average of 2,154 prompt tokens plus 723 extra-observation tokens, about 2,877 total against an 8,400-token text context, roughly 2.9x compression; for MRCR's longest bin the text reference is 197,909 tokens and the prompt 53,181 tokens, about 3.7x compression with observations; on 64K-128K four-needle MRCR, online latency falls from 187.09 s to 67.06 s.

Perspective

This work targets settings that require retrieval and reasoning over long documents, multi-document QA, and extended interaction histories, and applies to multimodal models with a visual encoder and tool-calling ability; its setting is a 9B-parameter backbone, 29.4K REL-CoT supervision examples, a default 72 DPI evaluation resolution, and a 144 DPI enhancement source. For a reader, it means long-context systems can retain global coverage with fewer input tokens and spend detail only where reasoning needs it; the paper also lists extending to larger backbones, more diverse documents, and larger context scales as future directions.

The paper states that the current study is limited in model and training-data scale, that the benefits of scaling model capacity and localization supervision remain to be established, and that long-context compression must still accommodate more languages, layouts, and reasoning demands. Ablations show gains concentrate on QA and synthetic retrieval where a few passages determine the answer, summarization changes little, few-shot results are mixed, and VTCBench Reasoning remains below the fixed-resolution backbone, so which task shapes benefit most from adaptive cropping stays an open question. In addition, this evidence bundle is parsed full text with figures rendered as text and tables, so checking figure-level detail still requires the original.

Sources