SRSP trains a style planner with a frozen TTS model's speech-likelihood reward, beating captioning baselines that see the target audio on five acoustic metrics
Synopsis
The work first shows that speech-text alignment only weakly predicts downstream acoustic similarity, then proposes Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner with GRPO using the teacher-forced likelihood of target speech tokens under a frozen TTS model as reward; on an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP outperforms the Base LLM and target-audio-informed captioning baselines on five downstream acoustic metrics and gains in LLM-judged contextual appropriateness and reference consistency.
Figure 1: Mean within-utterance Spearman correlations across both development partitions.
arXivInterpretation
In a candidate-level analysis, speech-text alignment (T2S) only weakly correlates with downstream acoustic similarity (S2S, emotion similarity, MCD-DTW), and selecting instructions by T2S improves only slightly over random selection. Prior work uses speech captions as pseudo-labels or intermediate representations, implicitly assuming descriptive fidelity stands for control effectiveness; this work tests that assumption directly with within-utterance Spearman correlations and selection experiments over eight candidate instructions per utterance. Eight candidate instructions are sampled per utterance on a 5,500-utterance development set and synthesized; T2S is computed with ParaCLAP-C, ParaCLAP-S, and CLSP and correlated within utterance against S2S, emotion2vec emotion similarity, and MCD-DTW. Table 1 shows T2S-based selection beats random only slightly, while selection by any downstream acoustic metric beats all three T2S criteria on all five metrics.
SRSP builds the reward from the frozen TTS model's mean teacher-forced cross-entropy loss over target speech tokens and updates the text-based style planner with GRPO, without requiring target style captions. Unlike supervision from speech descriptions, the reward comes from the downstream synthesizer's likelihood of the same target speech, so the planner is optimized for synthesis behavior rather than description agreement; auxiliary penalties discourage overly short, generic, and repetitive instructions. The planner is Qwen3.5-9B adapted with rank-16 LoRA, the frozen reward and synthesis model is Fun-CosyVoice3-0.5B, training runs one epoch with TRL 1.3.0 and DAPO-style loss normalization, and the clipped objective includes a KL penalty after group-relative reward standardization.
On the English ISCSLP 2026 CoT-TTS subset, SRSP achieves the best results on all five downstream acoustic metrics in both evaluation partitions, including against captioning baselines with access to the ground-truth target waveform. The unadapted Base LLM scores lower S2S and higher MCD-DTW than Raw TTS, showing that adding style instructions alone does not guarantee improvement; speech-rewarded post-training reverses this degradation. Four development/test partitions each contain 2,750 utterances (about 5 hours); dev_in/test_in hold out complete scenes from movies seen in training, dev_out/test_out hold out entire movies. AF-Next has the highest T2S under all three encoders yet ranks last on all five downstream metrics, mirroring the candidate-level mismatch.
In blind A/B judgments by Gemini 3.8 Flash, SRSP is preferred over all four baselines in both the context-only and target-reference settings. The context-only setting provides no target recording, so win rates against the pre-RL Base LLM (54.19%/53.67%) indicate the gain extends to contextual appropriateness rather than only fitting the reference. Blind A/B comparisons cover all utterances in both test partitions with two-sided exact binomial tests reported as significant; against ground-truth speech SRSP has the highest win rate among synthesized systems (18.41%/19.06%), while ground-truth speech is still preferred overall.
Perspective
The result targets research and engineering settings where natural-language instructions control conversational TTS style: at inference the planner receives only dialogue history and response text, emits one English style sentence, and conditions the TTS model together with the response text and a same-speaker prompt audio. The method applies to training stages where target speech is available to construct rewards, and the reward comes from the frozen synthesizer's teacher-forced likelihood, so it applies most directly to deployment with the same TTS backend used in training (here Fun-CosyVoice3-0.5B). The candidate-level finding applies to ranking candidate instructions for the same utterance, indicating that picking instructions by descriptive alignment within a candidate pool is unreliable.
The expressive-speech evaluation relies on Gemini 3.8 Flash as an automatic judge and has not been validated by human listening tests, so the contextual-appropriateness and reference-consistency win rates should be read as results under that judging setup. Reward computation and synthesis share CosyVoice3, and transfer to other TTS backends remains an open question. Ground-truth speech is still preferred overall in the context-only setting, indicating a remaining gap between synthesized systems and the target. In addition, equations and some numeric values appear as placeholders in the loaded text, so specific hyperparameters and loss details cannot be fully verified from it.
