Skip to main content
Back to timeline
arXivSource publication:

Fudan team rewards Chinese glyph structure via IDS decomposition, letting Qwen-Image lead on both structural quality and semantic alignment on LongText and GenTextEval

Synopsis

The work introduces IDSpect: it recursively decomposes target Chinese characters into Ideographic Description Sequence (IDS) tokens using the Unicode 16.0 BabelStone lexicon, trains an expert IDS recognizer built on an SVTRv2 backbone with an NRTR-style decoder to transcribe rendered text regions directly into IDS sequences, and fuses a token-level F1 with globally unique token credit against a whole-character semantic reward, so that GRPO post-training of Qwen-Image improves structural quality and semantic alignment on LongText and GenTextEval without changing the image generator or adding inference-time cost.

AI-generated editorial illustration: Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

Interpretation

The paper argues that OCR-based reinforcement-learning rewards compare decoded transcripts with target strings and treat each ideograph as an atomic character, so visually different radical-level errors receive equally coarse feedback, and linguistic priors can even recover the intended character from a malformed glyph, producing a high semantic score without adequate visual evidence. Earlier structure-aware evaluators such as TextPecker can flag anomalous glyphs but replace incorrectly written characters with '#', so once a glyph is collapsed into an anomaly marker its correct and incorrect internal structures are no longer distinguished; this work moves reward granularity from the whole character down to components and spatial operators. The paper illustrates both failure modes in the same generated image: a specialist OCR model follows local appearance and predicts a visually plausible but incorrect character contributing 0.00 to the character-level reward, a global MLLM OCR recovers the target from surrounding context despite its malformed structure producing a false-positive local contribution of 1.00, and TextPecker correctly identifies the glyph as structurally anomalous but still assigns 0.00, whereas IDSpect retains credit for matched IDS tokens while penalizing mismatched ones for a standalone IDS reward of 0.22.

The authors build an expert IDS recognizer that couples an SVTRv2 visual backbone with an NRTR-style Transformer decoder to autoregressively predict over the IDS vocabulary, mapping a text-region image directly to a sequence of spatial operators and reusable components rather than first recognizing a character sequence and then decomposing it symbolically. Prior Chinese OCR work uses radicals, strokes, and IDS to recognize rare or unseen characters; this work instead trains an IDS reader as a reward model and uses its structured output to optimize a generative model, noting that reward utility depends not only on final recognition accuracy but also on whether intermediate scores rank partially correct and malformed glyphs in a useful order. Training uses IDSynth-1M, one million rendered scene text images whose characters are sampled approximately uniformly from the character set covered by the selected Chinese fonts to suppress natural lexical co-occurrence and linguistic shortcuts; recursive decomposition depth is capped at 10 and the maximum decoded IDS sequence length is 100 tokens.

IDSpect aligns each detected text crop with candidate spans of the target IDS sequence through semi-global edit alignment, non-maximum suppression to remove near-duplicate matches, and a global assignment that lets each target position receive exact-match credit only once, making the comparison robust to the order of detected text regions, and uses a token-level F1 as the compositional reward added with equal weight to the whole-character semantic reward. Direct concatenation compares crop-level predictions with the target IDS sequence according to detector order and is sensitive to order mismatches, while character-wise matching reduces order dependence but treats characters independently and discards inter-character ordering information; crop-wise alignment matches each crop to a target span while preserving IDS-token order within each crop. The ablation shows crop-wise alignment achieves the highest Sem. of 0.911, improving over direct concatenation and character-wise matching by 0.025 and 0.012 respectively, with Qua. 0.979 only 0.001 below per-crop consistency; per-crop consistency reaches the highest Qua. of 0.980 but its Sem. of 0.901 is lower than the 0.911 of target-based crop-wise alignment.

After post-training Qwen-Image with Flow-GRPO under the Flow-Factory framework using LoRA adapters, IDSpect improves both structural quality and semantic alignment on LongText and GenTextEval, outperforming general OCR-based rewards and TextPecker. Relative to the OCR reward it exceeds it by 0.026 and 0.037 on GenTextEval for Qua. and Sem., and improves over the TextPecker reward by 0.006 and 0.014; on LongText it achieves the highest Qua. of 0.975 and Sem. of 0.928 while maintaining an Avg. of 0.972 comparable to TextPecker. Baselines include the frozen model, OCR-only GRPO, and TextPecker-guided GRPO, reported on LongText-Bench and GenTextEval-Bench following TextPecker's Chinese evaluation protocol with the original benchmark text score, structural quality, and semantic alignment; qualitative comparison covers three prompts with multiple Chinese text regions of different lengths and spatial arrangements, with the improvement especially apparent in longer lines.

Perspective

The result targets text-to-image post-training where Chinese ideographs are the rendering objective: during GRPO, a weighted sum of the IDS structural reward and the whole-character semantic reward guides Qwen-Image, and both reward branches are discarded after training, so the image generator is unchanged and no inference-time cost is added. It suits settings that need to assess whether components and spatial relations are faithfully rendered, such as signs and posters containing multiple text regions of different lengths and spatial arrangements. In the conclusion the authors list extending the framework to broader Chinese character coverage and multilingual scripts such as Japanese and Korean as future work, indicating that the current setting assumes the covered character set and selected fonts.

With naturally distributed text the IDS recognizer can first recognize a glyph as a frequent character and then reproduce its canonical IDS, reducing IDS prediction to character recognition; the paper uses approximately uniform sampling and suppressed lexical co-occurrence in synthetic data to discourage this shortcut, but how this design behaves on real long-tailed glyphs remains a question a reader would watch. Per-crop consistency reaches a higher Qua. of 0.980 than the 0.979 of the target-based reference while its Sem. of 0.901 is lower than 0.911, suggesting room to weigh structural consistency against target alignment. Evaluation also follows TextPecker's Chinese protocol and includes published Qwen-Image, OCR-reward, and TextPecker results for direct comparison, so readers concerned about comparability across evaluation protocols may treat this as a direction for further observation.

Sources