Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Kana-domain ASR rewards for Japanese TTS reinforcement-learning post-training cut target-kanji reading error by roughly 40% relative and converge faster

Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.