Skip to main content
Back to timeline
Apple Machine Learning ResearchSource publication:

After compressing a speech encoder 2.8x, the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs

Synopsis

This work studies how to compress, by distillation, the streaming neural audio encoder (tokenizer) used by on-device System-wide Dictation on Apple devices, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes; only the student encoder is trained to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch, and because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces and applies both to a tokenizer pretrained alone and to one jointly trained with a language model; at 2.8x compression the distilled student stays within 1.

AI-generated editorial illustration: Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Interpretation

The paper proposes using the pre-quantizer latent as the distillation supervision target rather than the discrete token or the output distribution, because it is the last representation the two token interfaces share. Relative to common distillation practice that targets discrete tokens or output distributions, this work chooses the pre-quantizer latent the model actually consumes, placing supervision before the quantizer and the language-model bridge. The paper presents this design as a method description, stating that only the student encoder is trained to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch.

Because the supervision target precedes both the quantizer and the language-model bridge, one distillation recipe covers both token interfaces the paper supports and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. This lets a single recipe serve different interfaces and different training pipelines without separate designs, broadening the method's applicability. The paper argues by design that the target precedes the quantizer and the language-model bridge, so one recipe covers both interfaces and both pretraining situations.

At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative. This gives a concrete degree of WER retention after compression and a relative gap against an identically sized independently trained baseline. The paper reports specific numbers: 2.8x compression, five of six teacher-student pairs, 1.9% relative WER, and 3.9% relative improvement, but the given text does not state the evaluation dataset, sample size, or statistical tests.

The compression targets the on-device dictation setting where an always-on tokenizer competes for the same DRAM as a sparsely activated foundation model, and its parameter count bears directly on power and latency. The paper grounds the motivation for compression in on-device memory competition, power, and latency rather than in model-size metrics alone. The paper states that the foundation model is sparsely activated under Instruction-Following Pruning, with only a small subset of experts occupying DRAM at any time, so the always-on tokenizer competes for the same memory.

Perspective

This work targets compression of the streaming neural audio encoder used by on-device System-wide Dictation on Apple devices, in settings where a tokenizer must be always-on alongside a sparsely activated foundation model; the two token interfaces the paper supports, and both the pretrained-alone and jointly-trained-with-a-language-model cases, fall within the same recipe. For practitioners seeking to reduce on-device memory, power, and latency, this result offers a reference point for retaining WER at 2.8x compression.

The given text is an incomplete excerpt and does not include the evaluation dataset, sample size, statistical tests, the specific teacher-student pair that did not meet the target among the six, the exact measure of 2.8x compression, or measured power and latency results. Readers judging applicability to their own setting would still need this information, which is not presented in the text.

Sources