Skip to main content
Back to timeline
arXivSource publication:

Pruned CTC removes linear-in-vocabulary memory growth from large-vocabulary CTC training, and LLM-CTC stays within 7% relative WER of LLM-CE on GigaSpeech while recognizing 7–10× faster

Synopsis

The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.

AI-generated editorial illustration: Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Interpretation

Pruned CTC restricts CTC alignment computation to the target classes present in a batch plus blank while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients. Previously CTC materialized frame-by-vocabulary activations in memory, which is costly at large vocabularies; Pruned RNN-T restricts the time–label region while keeping the vocabulary dimension unchanged, and Cut Cross-Entropy relies on a known target class per position, whereas CTC lacks frame-level targets and computes posteriors by forward–backward over the transcript. This work separates reduced-vocabulary alignment computation from full-vocabulary normalization and dense gradient reconstruction, using chunked recomputation to avoid materializing the full activation matrix. The paper gives a full proof of Proposition 1 in Appendix A, showing retained alignment probabilities are unchanged and gradients agree by the chain rule; Appendix B.2 runs 1,188 comparisons across 22 vocabularies (500 to 180,000 classes) and batch frame counts of 16,000/32,000/64,000, reporting maximum relative errors for loss, head-input gradients and weight gradients and identifying FP32 accumulation as the main error source.

With a Zipformer-M encoder and a 180K vocabulary, Pruned CTC reduces full-step memory by 5.1× with only 17% step-time overhead, and yields nearly identical WER/CER to standard CTC on LibriSpeech, GigaSpeech and AISHELL-1. The paper reports that head-and-loss activation memory no longer scales linearly with vocabulary size: at a 16k-frame batch it drops from 43.0 GiB to 1.29 GiB (33.3×), and standard CTC runs out of memory in some configurations where Pruned CTC completes all tested ones. Table 4 gives 2.3× to 5.1× full-step memory reductions and 14% to 17% step-time overhead across three batch sizes; Table 1 lists standard CTC and Pruned CTC numbers item by item on LibriSpeech test-clean/test-other, GigaSpeech dev/test and AISHELL-1 dev/test, with differences at the second decimal place.

Built on Pruned CTC, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies while retaining causal attention, staying within 7% relative WER of LLM-CE across six Qwen3 sizes with 7–10× faster recognition. Prior LLM-based ASR typically uses token-level cross-entropy with autoregressive generation (LLM-CE), and streaming often relies on chunk-level speech–text alignments; LLM-CTC adapts pretrained LLMs under utterance-level transcripts without chunk-level alignments. Table 3 lists dev/test WER and RTF for LLM-CE and LLM-CTC at six sizes from 0.6B to 32B on GigaSpeech: LLM-CTC test WER falls from 10.57% at 0.6B to 9.94% at 32B, with RTF values such as 0.0724 versus 0.0099; at 4B and above, LLM-CTC test WER is lower than the SPEAR-XLarge CTC and RNN-T baselines.

The streaming extension of LLM-CTC pairs utterance-level supervision with bounded-history inference and KV cache reuse, staying within 3% relative WER of matched offline models on the GigaSpeech test set. The streaming version avoids chunk-level speech–text alignments during training and produces each chunk's query emissions in one LLM forward pass; Appendix E.1 proves parallel batch computation and per-utterance cached computation are equivalent in exact arithmetic. Table 4 reports offline and 2/4/8-second left-history streaming WER for Qwen3-ASR 0.6B and 1.7B, with streaming increasing test WER by less than 3% relative to matched offline models; increasing left history from 2 to 8 seconds changes WER little.

Perspective

The results target speech recognition training with large-vocabulary CTC objectives: on the encoder side they are validated with a Zipformer-M encoder, vocabularies from 500 to 180,000 classes and batch frame counts of 16,000 to 64,000; on the LLM side with a SPEAR-XLarge v2 encoder paired with Qwen3 from 0.6B to 32B, and with LoRA fine-tuning of Qwen3-ASR 0.6B/1.7B, keeping the speech encoder and LLM head frozen. The streaming setting is bounded to 2-second chunks and 1 to 8 seconds of left history. The paper states code and pretrained models will be open-sourced, providing an entry point for reuse on other languages, encoders or CTC objectives.

The paper reports that Pruned CTC's head-and-loss activation memory no longer scales linearly with vocabulary size, but full-step reductions are limited by storage and activations outside the head and loss, giving 38.2% and 27.3% full-step reductions for 8B and 14B versus 24.8× to 25.7× at the head/loss level. Finite-beam alignment pruning can introduce discontinuities in the loss; the paper uses an alignment beam of 100 and supports that value in Appendix B.3 with numerical estimates that discarded posterior mass stays below a threshold, explicitly described as estimates rather than certified upper bounds. Appendix B.2 notes that when the two gradient terms nearly cancel for selected classes, rounding errors can dominate their difference, and gives precision strategies and error magnitudes. On streaming, increasing left history from 2 to 8 seconds changes WER little, which the paper describes as limited benefit from longer history. In addition, this evidence bundle is full text without the figure images themselves, so the specific curve shapes of Figures 4, 5, 7 and 8 can only be understood from the prose descriptions.

Sources