Skip to main content
Back to timeline
arXivSource publication:

Q-SPT compresses speech tokens with learnable queries, achieving the best reconstruction at a matched frame rate and leading downstream recognition and TTS perceptual quality

Synopsis

The work proposes Q-SPT, a low-frame-rate dual-stream speech tokenizer that uses two separately learned, context-aware query-based compressors for the semantic and acoustic streams, with an autoregressive text loss explicitly supervising the semantic compressor; experiments report the best reconstruction among the evaluated codecs at the same frame rate and, in downstream speech language models, the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.

Source-provided article image: Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Figure 1 ·

Figure 1: Comparison of temporal compression strategies for low-frame-rate speech tokenization: (a) average pooling in DualCodec [ 22 ] , (b) similarity-based merging in FlexiCodec [ 23 ] , and (c) learnable query-based compression in Q-SPT. (d) Detailed architecture and training objectives of Q-SPT. Dashed elements indicate components and connections used only during training.

arXiv

Interpretation

It introduces a dual-stream compression architecture in which fixed-rate queries independently attend to the semantic and acoustic streams, making compression stream-specific and context-aware. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, while similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream; Q-SPT instead uses learnable, stream-separated compressors. At the abstract level, the architecture is described and comparative experimental conclusions are reported, without dataset details, frame-rate values, or ablations.

It adds an autoregressive text loss that explicitly supervises the semantic compressor to preserve linguistic information at low frame rates. Linguistic preservation shifts from passively depending on compression rules to being directly constrained by a text-side objective on the semantic branch. The abstract states the loss design and its role, but gives no loss weights, training configuration, or isolated ablation results.

At the same frame rate, Q-SPT achieves the best reconstruction among the evaluated codecs and, in downstream speech language models, the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility. Compared with existing rule-based compression schemes, it reports leading performance on both reconstruction and downstream task measures. The conclusions are qualified by 'among the evaluated codecs' and 'at the same frame rate'; they are abstract-level comparative statements without metric values or statistical tests in the text.

Perspective

The results target settings where neural speech codecs serve as tokenizers for speech language models, especially where both linguistic information and acoustic detail must be preserved at low frame rates. The intended audience is researchers and engineering teams building or deploying speech language models who care about computational and memory costs; the value lies in turning compression from fixed rules into learnable, stream-separated aggregation, enabling comparison of reconstruction and downstream performance at the same frame rate.

Based only on the abstract, the specific frame rate, dataset composition, baseline choices, and metric values cannot be judged, nor can the individual contributions of the dual-stream separation, query-based aggregation, and text loss be confirmed. Both the best reconstruction and the best downstream results are qualified by 'among the evaluated codecs' and 'at the same frame rate'; their stability across languages, speaker conditions, and noisy environments remains to be observed, and whether a trade-off exists between the text loss and acoustic fidelity requires the original experiments.

Sources