Skip to main content
Back to timeline
arXivSource publication:

Tacit-TTS swaps autoregressive decoding for masked prediction to make zero-shot voice cloning over 10x faster, staying usable with eight cross-lingual references plus infant babble and synthetic gibberish

Synopsis

Tacit-TTS, distilled from IndexTTS2, replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, adds training-free acoustic length estimation, and accelerates the flow-matching renderer via ReFlow distillation, achieving competitive zero-shot cloning quality on two English and two Mandarin datasets while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds and supporting cross-lingual and non-lexical references without a reference transcript.

AI-generated editorial illustration: Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning

Interpretation

Tacit-TTS replaces IndexTTS2's autoregressive text-to-semantic stage with a masked non-autoregressive generator, accelerating the T2S stage by 26.5x, and cuts the flow-matching renderer from roughly 25 Euler steps to 4-8 steps through ReFlow distillation. Earlier non-autoregressive zero-shot TTS such as MaskGCT is typically trained on text-audio pairs, whereas this work distills from an autoregressive teacher and retains the teacher's speaker and emotion conditioning interface. The paper reports T2S latency dropping from 5274 ms to 199 ms and S2A from 875 ms to 310 ms, with total generation latency falling from 6149 ms to 509 ms, a 12.1x speedup; ablations select 12 T2S steps and 8 S2Mel steps as the default.

The work replaces dependence on reference transcripts or learned duration predictors with training-free acoustic length estimation: a reference pace factor multiplied by the target text's syllable count, with the pace clamped. Existing non-autoregressive systems mostly rely on reference transcripts or learned duration models, while this estimator computes directly from the reference waveform and target text, needing neither transcript nor training. Duration fidelity is measured by Pearson correlation between generated and ground-truth durations, reported as comparable to other systems, and the estimator is validated in cross-lingual and non-lexical settings.

On four zero-shot test sets, Tacit-TTS attains the highest speaker similarity among evaluated non-teacher baselines on the English datasets, content accuracy competitive with the strongest baselines, and perceptual quality close to IndexTTS2. The model is trained on far less data than the 55k hours used by IndexTTS2 yet retains much of the teacher's zero-shot cloning ability, which the authors partly attribute to a loss design transferring both discrete semantic targets and continuous latent features. On LibriSpeech test-clean SS 0.875 and WER 5.54; on SeedTTS test-en SS 0.849 and WER 2.34; on SeedTTS test-zh SS 0.829 and CER 1.71; on AISHELL-1 test SS 0.804 and CER 2.21; UTMOS and DNSMOS stay close to the teacher across all four test sets.

Transcript-free conditioning remains effective with cross-lingual and non-lexical references: with FLEURS references in eight languages, Tacit-TTS reaches SS 0.789 with WER 0.10 on English targets and SS 0.692 with CER 1.90 on Chinese targets, with error rates below 1 across all eight languages, and it exceeds IndexTTS2 in speaker similarity on infant babble. Transcript-dependent baselines degrade or fail in these settings because ASR transcripts are unreliable, for example MaskGCT failing on 33.3% or even 100% of cases for some languages, while both transcript-free systems generate for all references. The cross-lingual study uses three 6-12 second references from different speakers per language, each generating 10 English and 10 Chinese target sentences; the non-lexical study uses 3 references selected from 70 infant-babble recordings plus 10 synthetic-gibberish references.

Perspective

The result targets zero-shot TTS settings where a voice must be cloned from a short reference recording and a reference transcript is unavailable or unreliable, including cross-lingual references and non-lexical vocalizations; the distillation route suits teams that already have a strong autoregressive teacher and want to retain its capability with a smaller model and fewer inference steps. The paper states its goal is not to train a stronger model from scratch but to preserve as much of the teacher's performance as possible under a much smaller distillation budget, so the setting presupposes an available teacher system and publicly released research datasets.

The paper notes weaker Mandarin results, possibly reflecting imbalanced speaker coverage between English and Chinese; perceptual quality comes from no-reference predictors such as UTMOS and DNSMOS, which the authors caution may favor speech distributions similar to their training data and should be treated as supportive metrics. The speedup varies with utterance length, peaking around 30 seconds and narrowing past a minute as full self-attention in the non-autoregressive stages catches up. Cross-lingual and non-lexical experiments use limited reference counts (three per language, three infant-babble references, ten synthetic-gibberish references), so the scope of those conclusions awaits larger-scale validation. The ethics statement notes that transcript-free cloning lowers the barrier to cloning a voice from arbitrary recordings and recommends requiring consent from the cloned speaker plus safeguards such as audio watermarking and synthetic-speech detection.

Sources