Skip to main content
Back to timeline
arXivSource publication:

SP-DocReader cuts Qwen3-VL-4B Vary-600K character error rate from 2.40 to 1.11 with difference-aware self-play

Synopsis

SP-DocReader is a self-play framework for document OCR that aligns reference and generated readings with a longest common subsequence, applies contrastive relative scoring and direct negative log-likelihood supervision only at unmatched positions, and trains only the OCR module while the visual encoder and language backbone stay frozen; across Qwen3-VL-4B and InternVL-3.5-4B, three self-play rounds reduce Vary-600K character error rate relative to two-epoch supervised fine-tuning and improve DocBank, IIT-CDIP, DocVQA and InfographicVQA scores.

Source-provided article image: SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
Figure 1 ·

Figure 1 : Overview of the SP-DocReader training framework. A frozen opponent generates a synthetic reading of the document, and Reading Discrepancy Masking aligns it with the ground truth to isolate the errors. The main player then optimizes a dual objective that combines the Reading Discrepancy Loss for difference-aware correction with the Focused Fidelity Loss for direct supervision on the hard tokens, after which the opponent is refreshed.

arXiv

Interpretation

The paper formulates a discrepancy-focused self-play objective: Reading Discrepancy Masking aligns reference and generated token sequences by longest common subsequence, selects unmatched positions in both readings, and scores those positions while retaining each sequence's full conditioning prefix. Unlike full-sequence supervised fine-tuning, which supervises every target position, this method concentrates the training signal on tokens that the generated reading actually gets wrong, without treating the extracted fragments as new sequences for scoring. Section 3.2 gives the formal definitions, including the LCS alignment operator, the unmatched position sets and the masked log-likelihood score, and Appendix C provides the dynamic-programming recurrence and EOS handling.

Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions and is combined with the contrastive Reading Discrepancy Loss in the final objective. The paper notes that a purely relative objective can decrease by lowering the generated reading's score rather than raising the reference score, so direct supervision is needed to anchor the correction direction; the authors also derive the combined gradient to separate relative score optimization from direct supervision. In the ablation, removing the contrastive term raises Vary-600K CER from 1.11 to 1.88, and removing FFL lowers DocVQA ANLS from 84.89 to 82.96; both reduced variants still beat SFT-2, and combining the losses gives the best score in every column.

On both Qwen3-VL-4B and InternVL-3.5-4B, three self-play rounds improve in-distribution transcription and zero-shot document benchmarks relative to two-epoch supervised fine-tuning. The gains extend from the training domain to DocBank and IIT-CDIP without target-benchmark adaptation, and DocVQA improves on both backbones, indicating that better page reading accompanies stronger document question answering. On Qwen3-VL-4B, Vary-600K CER falls from 2.40 to 1.11, DocBank and IIT-CDIP CER drop by 1.90 and 2.29 percentage points, and DocVQA and InfographicVQA ANLS rise by 3.67 and 3.60 points; on InternVL-3.5-4B all three reading benchmarks improve while InfographicVQA ANLS moves slightly from 48.15 to 48.06.

Modular OCR-only tuning outperforms full-parameter tuning on zero-shot reading and training cost, and shows higher data efficiency when labeled pages are limited. Full tuning achieves the lowest in-domain Vary-600K CER (0.98 versus 1.11), but its DocBank and IIT-CDIP CER are 11.28 and 16.87, higher than the modular variant's 9.05 and 14.42; modular tuning uses 0.5B trainable parameters versus 3.3B for OCR-free and 3.8B for full-parameter tuning. At a 30k-page budget, SP-DocReader reaches a five-task AVG of 82.46, above the 80.73 obtained by SFT with 120k pages; at the third round, modular tuning reaches a four-task zero-shot AVG of 79.58 at a relative training cost of 0.33, compared with 74.98 at 1.58 for full tuning.

Perspective

The result targets single-page transcription in settings with paired images and reference text, where freezing the visual encoder and language backbone and updating only the OCR branch is acceptable; the paper validates two vision language backbones on English and Chinese training pages and reports lower CER than SFT at all four tested resolutions from 48 to 150 DPI. For teams seeking to reduce labeled-page requirements while retaining zero-shot reading and document question answering, the modular recipe offers a reusable training procedure and cost accounting.

The paper states that it focuses on single-page transcription rather than multi-page reading or structured extraction, that its fixed generation budget may limit coverage of longer pages, and that the results do not establish generalization to other languages or model families. Appendix E notes that improvements are not monotonic across all metrics, for example Qwen3-VL-4B's Vary-600K NED moves from 0.988 at SP-DR-2 to 0.987 at SP-DR-3, and the main comparison fixes the third-round endpoint. Appendix C notes that masking does not automatically reduce projection computation, that a matched position need not have zero gradient in the internal computation, and that cost comparisons must cover preprocessing, generation, alignment and both scorers; these are open questions for readers assessing reproduction and deployment.

Sources