Public articles linked to the same research event.
arXiv Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.
Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.
Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.
Motivated by the mismatch between Japanese orthography and pronunciation, this work replaces orthographic CER with a kana-domain CER (Kana-CER) computed from a kana-transcribing ASR model and reference readings, and uses it as the reward for GRPO post-training of Sarashina2.2-TTS; on the Joyo Kanji Yomi Benchmark, the Kana-CER reward reduces target-kanji reading error by roughly 40% relative to the orthographic CER reward while keeping comparable orthographic CER, speaker similarity, and objective speech quality, and reaches its best validation performance in about 4k steps versus 18k steps, while unregularized Kana-CER optimization causes severe output elongation that KL regularization substantially suppresses.