Public articles linked to the same research event.
arXiv The work trains 15 masked-diffusion codec TTS models (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweeps refinement steps T in [1,16] at inference, measuring zero-shot synthesis on 174 held-out speakers via ASR word error rate and speaker verification; refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range (a 1.86x asymmetry), Best-of-K search recovers identity with 64.6-79.0% win rates across four independent encoders, a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3), and 62% of the remaining identity deficit lies in the codec rather than the generator.
The work trains 15 masked-diffusion codec TTS models (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweeps refinement steps T in [1,16] at inference, measuring zero-shot synthesis on 174 held-out speakers via ASR word error rate and speaker verification; refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range (a 1.86x asymmetry), Best-of-K search recovers identity with 64.6-79.0% win rates across four independent encoders, a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3), and 62% of the remaining identity deficit lies in the codec rather than the generator.
The work trains 15 masked-diffusion codec TTS models (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweeps refinement steps T in [1,16] at inference, measuring zero-shot synthesis on 174 held-out speakers via ASR word error rate and speaker verification; refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range (a 1.86x asymmetry), Best-of-K search recovers identity with 64.6-79.0% win rates across four independent encoders, a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3), and 62% of the remaining identity deficit lies in the codec rather than the generator.
The work trains 15 masked-diffusion codec TTS models (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweeps refinement steps T in [1,16] at inference, measuring zero-shot synthesis on 174 held-out speakers via ASR word error rate and speaker verification; refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range (a 1.86x asymmetry), Best-of-K search recovers identity with 64.6-79.0% win rates across four independent encoders, a separable B(d)B(T) model fits significantly better than substitution models (Delta AICc=+69.3), and 62% of the remaining identity deficit lies in the codec rather than the generator.