Skip to main content
Back to timeline
arXivSource publication:

Kana-domain ASR rewards for Japanese TTS reinforcement-learning post-training cut target-kanji reading error by roughly 40% relative and converge faster

Related research and updates

Synopsis

Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.

Source-provided article image: Pronunciation-Oriented Reinforcement Learning for Japanese Text-to-Speech with Kana-Domain ASR Rewards
Figure 1 ·

Figure 1: Evaluation metrics during RL on the validation split with β = 0.04 \beta=0.04 . Curves show the mean across five generation seeds with ± 1 ​ σ \pm 1\sigma shading. Rings indicate the checkpoints selected by the common validation criterion. The pre-RL model is shown at step 0.

arXiv

Interpretation

Computing the ASR reward in the kana domain rather than the orthographic domain more directly optimizes context-dependent Japanese kanji readings. Prior TTS RL post-training generally used orthographic CER/WER as an intelligibility reward; this work is, to the authors' knowledge, the first to investigate kana-domain ASR rewards for pronunciation-oriented RL post-training of Japanese TTS, turning Kana-CER from an evaluation metric into a training reward. On the Joyo Kanji Yomi Benchmark, Kana-CER (kanji) under the Kana-Whisper evaluator drops from 9.65 with the orthographic CER reward to 7.17, a relative reduction of about 40%; an independent wav2vec2-hiragana evaluator shows the same trend, indicating the gain is not specific to the Kana-Whisper model used in the reward pipeline.

The pronunciation gain does not come at the cost of orthographic accuracy or speech naturalness. The two reward conditions yield comparable orthographic CER (4.10 vs. 4.15) and similar speaker similarity and objective quality on the Japanese subset of CV3-Eval, indicating that the difference in reward representation does not degrade other dimensions. Orthographic CER is computed with the same Whisper large-v3-turbo model; SIM-o, UTMOS-s, UTMOS-v2, P.808, and OVRL are similar across the pre-RL model and the two GRPO checkpoints, with main reward-comparison results averaged over five generation seeds.

The kana-domain reward yields more efficient pronunciation-oriented optimization. Under the common validation criterion, the Kana-CER reward reaches its best validation performance at about 4k steps versus about 18k steps for the orthographic CER reward, and it improves target-kanji reading accuracy earlier and maintains lower target-reading error over subsequent checkpoints. Checkpoints are evaluated every 1k steps and selected by the mean of validation orthographic CER and sentence-level Kana-CER; moreover, RL uses only roughly half of the synthetic pronunciation data used for Stage 2 SFT, yet the Kana-CER model achieves lower error than Stage 2 SFT on every metric in Table 1.

Unregularized Kana-CER optimization induces severe output elongation, which KL regularization substantially suppresses. The work explicitly reports this length-degeneration phenomenon under the kana-domain reward and identifies KL regularization as a mitigation, keeping the selected checkpoint close to the pre-RL model in output duration on the Joyo Kanji Yomi Benchmark. Without KL regularization, Kana-CER more than doubles the mean completion length within the first few thousand steps, whereas the orthographic CER reward causes only a moderate change; under the regularized setting used in the main experiments this elongation is substantially suppressed, with Dur./base of 0.985.

Perspective

The result targets Japanese TTS post-training settings where context-dependent kanji readings must be realized correctly, and applies to training pipelines that have reference kana readings and access to a kana-transcribing ASR model; for teams seeking pronunciation gains in fewer optimization steps while preserving orthographic accuracy and speech quality, the Kana-CER reward is a drop-in alternative to orthographic CER, to be paired with KL regularization to control output duration.

Several numbers are missing or replaced by placeholders in the text and tables, including the exact relative-reduction percentage, the number of kanji-reading pairs and distinct kanji covered by the RL data, the per-pair sentence cap, the learning rate, temperature, top-p, maximum completion length, completions per input, reward-transformation parameter, and KL coefficient, as well as the Kana-CER value Kana-Whisper achieves on JSUT recordings; verifiable quantitative details therefore require the original tables and figures. In addition, length degeneration is reported only for the unregularized setting, so the trade-off across KL strengths and behavior at other model scales and TTS architectures remain open questions.

Sources