TS-SP two-stage LoRA adaptation cuts Qwen2.5-Omni-7B equal error rate on Vox1-O from 7.01% to 4.37%
Related research and updatesSynopsis
The authors propose TS-SP (Two-Stage Speaker Preservation): first adapt the native audio encoder of Qwen2.5-Omni-7B with speaker identity supervision under LoRA, then freeze that encoder and train only the language model with LoRA for pairwise speaker comparison, keeping all pretrained base weights fixed; on Vox1-O this lowers equal error rate from 7.01% for the Paired Loss Adaptation Baseline to 4.37% and minDCF from 0.915 to 0.611, with Vox1-E and Vox1-H also improving from 7.64% to 4.41% and from 19.31% to 9.43%, equal error rate staying within 4.31–4.79% under unseen prompts, while cross-domain CN-Celeb equal error rate is comparable to the baseline (16.74% versus 16.49%) with lower accuracy at the native decision threshold.
Figure 1: TS-SP: Stage 1 learns audio LoRA with intermediate speaker supervision; Stage 2 freezes the adapted audio pathway and trains LM LoRA for speaker comparison. All pretrained base weights remain fixed.
arXivInterpretation
TS-SP places speaker-discriminative representation learning inside the ALLM's own audio encoder rather than adding an external speaker encoder, and inference uses the native audio-to-language pathway without the speaker head or an additional connector. Prior work either projects frozen external speaker representations such as ECAPA-TDNN or ReDimNet into the language-model input space, or applies a joint objective on QFormer features; TS-SP instead applies speaker supervision within the native audio encoder before a separate pairwise language-model adaptation. The method is specified: Stage 1 uses parameter-free statistics pooling plus an AAM-Softmax speaker head and updates only audio encoder layers 0–30 LoRA (3.809 M) and the 2.191 M speaker head; Stage 2 updates only the 10.093 M language-model LoRA, and both systems retain 14.025 M LoRA parameters at inference.
Across the Vox1-O/E/H trial lists, TS-SP consistently lowers equal error rate and raises overall accuracy relative to the same-backbone Paired Loss Adaptation Baseline. The baseline ports the pairwise audio question-answering paradigm to the same backbone without speaker pretraining; TS-SP lowers equal error rate from 7.01% to 4.37% on Vox1-O, from 7.64% to 4.41% on Vox1-E, and from 19.31% to 9.43% on Vox1-H, with Vox1-H nontarget accuracy rising from 52.77% to 75.42%. Results come from Table 1; training uses only original VoxCeleb2 development recordings without data augmentation or speed perturbation, the baseline and Stage 2 share 200,000 recording pairs (100,000 same-speaker, 100,000 different-speaker), and reported numbers are final checkpoints rather than checkpoints selected by validation performance.
Equal error rate stays stable across several prompt rewrites beyond the training prompt, but fixed-threshold decisions do not. P1–P5 are unseen and require no further fine-tuning, with equal error rate within 4.31–4.79%; numeric answers (P3) give the highest equal error rate and minDCF, and the Chinese prompt P5 keeps equal error rate close to P0 (4.50% versus 4.37%) while shifting target/nontarget accuracy to 99.88%/79.88%, indicating stronger same-speaker bias. Table 2 reports equal error rate, minDCF, and both accuracy classes for six prompts; the authors note that P2–P5 change both wording and verbalizers, so the effect of answer format cannot be isolated, and P5 additionally changes language.
On cross-domain CN-Celeb evaluation, the gains from two-stage speaker supervision do not carry over, and equal error rate is comparable to the baseline. Both fine-tuned systems improve equal error rate over zero-shot Qwen2.5-Omni-7B (31.93%), but TS-SP's 16.74% is slightly higher than the baseline's 16.49% and its minDCF only marginally lower (0.984 versus 0.985); overall accuracy at the native threshold drops from 89.72% zero-shot to 34.89% for TS-SP. Cross-domain evaluation randomly samples 10% of CN-Celeb trials while preserving the original target-to-nontarget ratio, with no additional fine-tuning; that list is 99.49% nontarget trials, and the authors note that always answering No would reach 99.49% overall accuracy, so the metric is dominated by nontarget decisions.
Perspective
The result targets researchers and engineering teams who need speaker identity cues retained inside an audio large language model, and it applies to backbones with accessible intermediate audio features and separately adaptable audio and language components, instantiated here on Qwen2.5-Omni-7B. Training uses original VoxCeleb2 development recordings without data augmentation or speed perturbation, audio at 16 kHz with at most the first 6 s per utterance, and pairs concatenated with 1 s of silence; inference uses the native audio-to-language pathway without a speaker head or extra connector. The reusable parts are the two-stage LoRA procedure and the answer-token-probability scoring, and the authors list cross-domain generalization and joint training to strengthen speaker identity modeling while preserving speech recognition, content understanding, and instruction following as next steps.
The authors note the results support TS-SP as a whole but do not isolate the contribution of intermediate speaker supervision from the effects of the training schedule and additional optimization; P2–P5 change both wording and verbalizers, so the effect of answer format cannot be separated, and P5 additionally changes language. The answer-token-probability score is not necessarily a calibrated speaker-hypothesis likelihood ratio, and similar equal error rates across prompts do not guarantee stable minDCF or fixed-threshold decisions. The CN-Celeb list is 99.49% nontarget trials, so overall accuracy is dominated by nontarget decisions and the zero-shot model's higher accuracy should not be read as stronger speaker discrimination. Whether speech recognition, content understanding, and instruction following are retained also remains to be evaluated, as do other backbones.
