NVAlign backpropagates a frozen NV-ASR reward through the flow-matching sampler, lifting human-rated tag-following on three continuous AR-FM TTS systems by up to 39.8% to 54.0%
Synopsis
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
Interpretation
The paper introduces NVAlign, which treats a frozen NV-ASR model's probability on target tag tokens as a differentiable reward and backpropagates it through the generation process of a continuous AR-FM TTS model, jointly updating the autoregressive backbone and the flow head. Continuous AR-FM TTS previously lacked an established post-training method for non-verbal control: discrete-token models admit likelihood-ratio methods such as GRPO, while continuous sampling has no token probabilities; Flow-GRPO adapts likelihood-ratio optimization via stochastic differential equations, whereas NVAlign takes the direct-gradient route. The method is fully formulated: a two-step gradient surrogate replaces full-trajectory backpropagation with two DiT flow-head evaluations and routes gradients to the backbone through a straight-through latent connector; experiments compare against Flow-GRPO from the same SFT checkpoint on VoxCPM2, dots.tts, and an English production system.
On separate blinded listening sets, NVAlign raises tag-following accuracy over SFT across the board: VoxCPM2 from 43.2% to 47.6%, dots.tts from 63.2% to 65.4%, and the English production system from 39.8% to 54.0%. This extends direct-gradient reward optimization from image diffusion and from speech objectives such as intelligibility or speaker identity to non-verbal tag-following in continuous AR-FM TTS, with same-direction human results across three systems. Each language uses a 100-text listening set in which three blinded raters judge the target sound at its marked position and all three must agree; paired 95% bootstrap confidence intervals exclude zero for both English gains over SFT, both Mandarin VoxCPM2 gains over Flow-GRPO, and the English 100-text gain over Flow-GRPO, while intervals for all other comparisons include zero.
Ablations show the NV-ASR reward alone does not capture output quality, and that fidelity penalties plus reference-velocity regularization are what turn a higher reward into audible sound. The paper quantifies reward gaming: removing velocity regularization raises the reward from 0.468 to 0.488 while Gemini accuracy/PE drops from 3.54/3.34 to 2.88/1.94; removing the UTMOS term next lowers its predicted quality score from 3.39 to 3.01; removing the speaker terms then sharply reduces SIM from 0.796 to 0.360. Cumulative ablations remove constraints one at a time under a unified setting while reporting reward, Gemini ratings, UTMOS, and SIM together; the paper notes SIM and UTMOS are part of the training objective, so it also reports DNSMOS and human naturalness MOS, the latter staying within 0.1 of SFT on all three systems.
NVAlign largely preserves speech quality and speaker similarity while improving tag-following, but tag presence, expressive quality, and speech quality need not move together. The paper separates three metric families: relative to SFT, SIM, UTMOS, and DNSMOS rise for all three systems; Mandarin CER rises by 1.26 and 0.13 points, and English WER rises by 0.07 on the listening set and 1.33 on NVV-SuperBench; on the English benchmark human tag-following rises from 33.7% to 53.7% while Gemini perceptual effect decreases. Metrics cover benchmark outputs and 100 held-out clips per system, with CER measured by Paraformer-zh and WER by Whisper-large-v3 on reference transcripts without NVV tags; the paper uses this to argue tag presence, expressive quality, and speech quality should be evaluated separately.
Perspective
The framework targets non-verbal tag-following in continuous autoregressive flow-matching TTS, in settings where NVV-annotated speech exists and an NV-ASR reward model can be trained; the paper validates it on VoxCPM2, dots.tts, and an English production system, covering a 39-tag inventory and two languages. For practitioners it supplies a post-training recipe: supervised-fine-tune the TTS and NV-ASR models, freeze the recognizer, backpropagate the reward through a two-step gradient surrogate, and apply fidelity penalties and reference-velocity regularization at the same time. The formulation assumes discrete non-verbal events, so continuous or overlapping vocalizations fall outside the current scope, and the reward model itself is limited by scarce examples of rare vocalizations.
The reward comes from a learned NV-ASR model, and optimizing it can exploit recognition shortcuts or degrade properties such as prosody, because tag likelihood does not capture the full speech distribution; fidelity penalties and reference-velocity regularization constrain this drift, but adding objectives makes gradient balance itself part of the optimization problem. Statistical support for the human gains is uneven: confidence intervals exclude zero for both English gains over SFT, both VoxCPM2 gains over Flow-GRPO, and the English 100-text gain over Flow-GRPO, while other comparisons include zero, and dots.tts is 2.7 points below SFT on the NVV-SuperBench human subset. In addition, SIM and UTMOS participate in the training objective, so their gains should be read alongside DNSMOS and human naturalness MOS. The current formulation assumes discrete non-verbal events, leaving continuous or overlapping vocalizations to be examined.
