Skip to main content
Back to timeline
arXivSource publication:

Script-Aware Sparse Mixture-of-Experts: One Model Reading Ten Scripts

Synopsis

The work builds TextMuSS-10M, a synthetic scene-text dataset spanning 10 scripts and 229 languages, together with the real TextMuSS-Bench, and proposes ScriptMoE, a script-aware sparse Mixture-of-Experts recognizer that shares one visual encoder and uses an image-level router to activate Top-2 script-aligned experts plus an always-on shared expert, reaching 82.06% average accuracy on TextMuSS-Bench (1.31% above the strongest STR baseline) and lifting CC-OCR end-to-end multilingual F1 from 65.71% to 80.89% by replacing only the recognizer in PP-OCRv5.

AI-generated editorial illustration: All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Interpretation

It constructs the TextMuSS-10M synthetic dataset and the TextMuSS-Bench real benchmark, providing balanced supervision for scripts where real data is scarce. The paper notes SynthMLT was the only available multilingual synthetic STR dataset, but its scale and quality were too limited and its coverage was the ten MLT2019 languages; the authors synthesize 1M samples per script (10M total) with the UnionST engine and merge the ten per-script character tables into a unified vocabulary of 19,684 characters including space. The data ablation shows real-only training collapses to 0.00 on Russian, Thai and Tibetan, indicating missing real supervision; synthetic-only surpasses real-only on the MLT2019 average by 2.49%, and synthetic plus real exceeds real-only by 8.40%.

It proposes ScriptMoE: one shared visual encoder, an image-level router dispatching each image to Top-2 script-aligned experts, and an always-on shared expert carrying cross-script knowledge. The ten scripts are grouped into four expert groups by character morphology (alphabet Latin+Cyrillic, CJK, Arabic family, Others), and routing is decided once per image rather than per token; unlike incremental multilingual schemes such as MRN that keep a separate feature extractor per language with parameters growing linearly, the expert count here is fixed at four. Model ablations show the pure AR baseline without experts averages 80.75%, four experts reach 82.06%, and ten experts fall back to 81.32%; removing the shared expert drops to 81.35% (Arabic -3.62%, Tibetan -3.65%) and removing the script-classification signal drops to 81.63%.

It systematically evaluates all-in-one multilingual STR under a unified protocol, with ScriptMoE achieving the best average accuracy on TextMuSS-Bench and surpassing expert systems and strong VLMs end-to-end. Fifteen STR baselines and nine general OCR systems are retrained on the same data for the same epochs and re-evaluated, while generalist and OCR-specialized VLMs are evaluated zero-shot; ScriptMoE averages 82.06%, 1.31% above the strongest baseline SVTRv2-AR, with gains concentrated on Arabic (+2.98%), Thai (+2.40%) and Tibetan (+1.96%). On the CC-OCR end-to-end multilingual task with the PP-OCRv5 detector fixed and only the recognizer replaced, F1 rises from 65.71% to 80.89%, slightly above the strongest zero-shot general VLM Qwen3.5-9B at 80.73%, the OCR-specialized Qianfan-OCR at 76.70% and GoogleOCR at 71.78%; the paper states the latter has roughly a hundred times more parameters. Per forward pass, 41.13M of 45.85M parameters (89.69%) are activated.

It verifies that the router learns script-aligned expert specialization rather than an arbitrary partition. A four-way script-classification head shares the pooled router input but not the router weights and is trained with cross-entropy at weight 0.1; the paper visualizes router softmax and the resulting effective expert mixture for one representative line per script group. The hyper-parameter sweep shows over-weighting is equally harmful: at 0.2 and 0.5 the average falls to 81.44% and 81.35%, with Tibetan down to 84.55% at 0.5; a shared-expert width ratio of 0.5 already recovers nearly all the benefit, and enlarging it to 1.0 does not help.

Perspective

The result targets scene text recognition and end-to-end OCR deployment where one model must serve many scripts, especially long-tail scripts with scarce real data; the synthetic data and training protocol, the four expert groups, image-level Top-2 routing and the 0.1 script-classification weight are the preconditions for reproducing it. The paper also reports competitive results on Chinese BCTR (86.85% average, 0.92% above SVTRv2-AR) and English Union14M-Benchmark (88.95% average, 0.28% above), indicating the multilingual gains do not come at the expense of high-resource scripts. The authors propose future directions including continual learning to add new scripts without retraining the full model and integration with stronger text detectors.

The paper itself lists three limitations: a domain gap remains between synthetic data and real imagery; Latin-Cyrillic homoglyph confusion is alleviated but not fully solved, with Russian at 61.20% under synthetic plus real versus 68.13% under synthetic only; and the end-to-end pipeline still depends on the upstream PP-OCRv5 detector, whose missed and false detections are most visible on Latin scripts under strict word-level evaluation. In addition, the newly collected Russian, Thai and Tibetan portions of TextMuSS-Bench were each annotated by a single expert for that script and then checked at the morphological level by a separate reviewer not specialized in those low-resource languages, so annotation consistency is worth watching; the paper also notes a cross-script trade-off in which optimizing aggregate accuracy sacrifices per-script peaks.

Sources