EmoRES splits emotion vectors into shared and residual parts, lifting emotion hit rate by up to 12.95 points on IndexTTS-2 and CosyVoice2
Synopsis
The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
Interpretation
The authors provide a functional decomposition of emotion steering vectors for LM-based TTS, finding that every categorical emotion vector contains a shared component pointing from the neutral mean toward the centroid of emotional activations and a residual component pointing from the centroid toward the requested emotion. CoCoEmo previously treated each emotion vector as an indivisible direction governed by a single global strength; this work is the first to examine the shared displacement and the request-dependent residual separately. Component ablations at fixed steering strength and vector norm show that shared-only steering drops rank correlation to 5.30% (IndexTTS-2) and 7.20% (CosyVoice2), while residual-only steering raises rank correlation but yields lower E-SIM and TEP than the combined direction; shared-only steering also reduces the recognizer's neutral posterior from 36.51% to 7.49%.
EmoRES assigns an independent coefficient to the shared and residual components and strengthens the residual relative to the shared component, without retraining the backbone or learning a new subspace. At a residual coefficient of 1 the method reduces exactly to CoCoEmo, so conventional steering becomes a special case of the proposed family rather than a distinct rule. Every steered condition is rescaled to the mean Euclidean norm of the original emotion vectors before injection, so conditions differ only in direction and component weighting, not in total steering magnitude.
On the held-out IEMOCAP test set, EmoRES beats CoCoEmo on both IndexTTS-2 and CosyVoice2: rank correlation improves by 26.13 and 12.97 percentage points (relative gains of 118.8% and 33.1%), and emotion hit rate by 12.95 and 6.92 points (relative gains of 20.1% and 9.8%). The gains concentrate in rank correlation and hit rate, the two metrics measuring whether posterior changes follow the requested ordering, while TEP and E-SIM also improve in every comparison. CREMA-D serves as the development set for tuning, while IEMOCAP is strictly held out and contributes no speakers or recordings; steering vectors are extracted once from 8,050 paired neutral-emotional utterances (about 12.6 hours of speech) drawn from ESD, CREMA-D and RAVDESS.
Human evaluation shows the most frequent listener annotation matching the dominant target emotion rising from 54.72% to 73.89% (IndexTTS-2) and from 58.18% to 77.88% (CosyVoice2), with fidelity also improving and naturalness preference scores of 63.80% and 60.32%. Fidelity measures agreement with the complete target emotion mixture, so the gain in dominant-emotion hit rate does not come at the cost of suppressing the remaining requested emotions. Each sample received at least 9 independent annotations from different annotators, with 360 IndexTTS-2 samples and 330 CosyVoice2 samples evaluated, and paired 95% confidence intervals excluding zero.
Perspective
The result targets speech synthesis research and engineering settings that want stronger emotion controllability without retraining a backbone: steering vectors are extracted once from paired corpora (ESD, CREMA-D, RAVDESS) and reused, injection happens at the attention output during inference, backbone parameters stay frozen, and no target-emotion reference recording is needed. Evaluation covers the IndexTTS-2 and CosyVoice2 backbones, the CREMA-D development set and the strictly held-out IEMOCAP test set, with emotions limited to angry, happy, sad and surprise. The authors state that future work will extend EmoRES to token-level or segment-level control for emotions that vary within an utterance and test whether the shared-residual decomposition generalizes across additional emotions, languages, speech corpora and TTS architectures.
The component ablations and random-direction controls show that performance depends on the measured directions of both components, but the functional interpretation that the shared component moves speech away from neutral and the residual directs it toward the requested emotion remains the authors' hypothesis, and the mechanistic causal chain awaits further testing. The evaluated emotion set is angry, happy, sad and surprise, mixed-emotion evaluation relies on corpus-annotated proportions, and emotions varying within an utterance are not yet covered. Sample and annotator counts for human evaluation are given in the text, but annotator demographics and language background are not elaborated here. This reading is of the full text, yet some figure and table details appear in the text by reference, so reproducing exact numbers still requires the original appendices.
