Skip to main content
Back to timeline
arXivSource publication:

Refinement buys intelligibility, search buys identity: depth and inference steps are not interchangeable in masked-diffusion TTS

Related research and updates

Synopsis

The work trains 15 masked-diffusion codec TTS models (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweeps refinement steps T in [1,16] at inference, measuring zero-shot synthesis on 174 held-out speakers via ASR word error rate and speaker verification; refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range (a 1.86x asymmetry), Best-of-K search recovers identity with 64.6-79.0% win rates across four independent encoders, a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3), and 62% of the remaining identity deficit lies in the codec rather than the generator.

Source-provided article image: Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS
Figure 3 ·

Figure 3 : Identity is mostly not the model’s to win. (A) what each axis buys in speaker similarity, against the headroom that remains to the codec ceiling; bars are coloured by which lever they are ( rented at serve time , owned in training , shape ). Refinement’s bar is the largest only because it is measured from one-step decoding, which is not an operating point anyone ships; from the deployable T = 16 {\color[rgb]{0.0586,0.4805,0.4453}T}{=}16 default it buys nothing further, and test-time search is what moves identity from there. (B) of the distance between real audio and our best system, 62% is lost in the codec itself.

arXiv

Interpretation

Refinement steps and model depth act unequally across TTS capabilities: refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range, a 1.86x asymmetry. Where inference budget and parameter count are often treated as interchangeable forms of compute, this work measures them separately per capability and reports an asymmetry robust across multiple error metrics. 15 models, 3 seeds, 2,000 hours of speech, 174 held-out speakers, a T in [1,16] sweep, and measured floors as reference; the asymmetry holds across multiple error metrics.

Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. It extends test-time compute beyond the single refinement axis to a search axis, showing search targets a different bottleneck than refinement. Win-rate ranges measured on four independent encoders, giving cross-encoder independent validation.

Depth and steps are not interchangeable: a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3). It supports the claim that depth and steps act independently through model comparison rather than curve shape alone. Model comparison based on Delta AICc, with a difference of +69.3.

62% of the remaining identity deficit lies in the codec rather than the generator; retraining at 3x and 6x schedules attenuates the gap from 1.84 to 1.36 to 1.23x without reversing it, because intelligibility saturates with steps while identity keeps improving. It localizes the identity bottleneck to the codec stage and shows that longer training schedules weaken but do not remove the asymmetry. Retraining experiments report the three ratios 1.84, 1.36, and 1.23, alongside the observation that intelligibility saturates while identity continues to improve.

Perspective

The result applies to the masked-diffusion codec TTS setting: training on 2,000 hours of speech, zero-shot evaluation on 174 held-out speakers, and an inference-step sweep of T in [1,16]. Within that scope it supports treating refinement steps and model depth as resources aimed at different bottlenecks, and it points to Best-of-K search for the identity dimension. For TTS system designers allocating inference budget, this means intelligibility can be pursued primarily by adding refinement steps, while identity gains depend more on search or on codec-side changes.

Readers should still watch: whether the asymmetry ratio holds outside T in [1,16]; how Best-of-K win rates change with K and how its extra inference cost is accounted for; under which codec configuration the 62% codec deficit was measured; and whether the attenuation trend at 3x and 6x schedules continues. This document is summary-level material without figures or full experimental detail, so the measurement conditions and ablation settings behind these numbers need to be checked against the original.

Sources