Public articles linked to the same research event.
arXiv The authors propose TS-SP (Two-Stage Speaker Preservation): first adapt the native audio encoder of Qwen2.5-Omni-7B with speaker identity supervision under LoRA, then freeze that encoder and train only the language model with LoRA for pairwise speaker comparison, keeping all pretrained base weights fixed; on Vox1-O this lowers equal error rate from 7.01% for the Paired Loss Adaptation Baseline to 4.37% and minDCF from 0.915 to 0.611, with Vox1-E and Vox1-H also improving from 7.64% to 4.41% and from 19.31% to 9.43%, equal error rate staying within 4.31–4.79% under unseen prompts, while cross-domain CN-Celeb equal error rate is comparable to the baseline (16.74% versus 16.49%) with lower accuracy at the native decision threshold.
The authors propose TS-SP (Two-Stage Speaker Preservation): first adapt the native audio encoder of Qwen2.5-Omni-7B with speaker identity supervision under LoRA, then freeze that encoder and train only the language model with LoRA for pairwise speaker comparison, keeping all pretrained base weights fixed; on Vox1-O this lowers equal error rate from 7.01% for the Paired Loss Adaptation Baseline to 4.37% and minDCF from 0.915 to 0.611, with Vox1-E and Vox1-H also improving from 7.64% to 4.41% and from 19.31% to 9.43%, equal error rate staying within 4.31–4.79% under unseen prompts, while cross-domain CN-Celeb equal error rate is comparable to the baseline (16.74% versus 16.49%) with lower accuracy at the native decision threshold.
The authors propose TS-SP (Two-Stage Speaker Preservation): first adapt the native audio encoder of Qwen2.5-Omni-7B with speaker identity supervision under LoRA, then freeze that encoder and train only the language model with LoRA for pairwise speaker comparison, keeping all pretrained base weights fixed; on Vox1-O this lowers equal error rate from 7.01% for the Paired Loss Adaptation Baseline to 4.37% and minDCF from 0.915 to 0.611, with Vox1-E and Vox1-H also improving from 7.64% to 4.41% and from 19.31% to 9.43%, equal error rate staying within 4.31–4.79% under unseen prompts, while cross-domain CN-Celeb equal error rate is comparable to the baseline (16.74% versus 16.49%) with lower accuracy at the native decision threshold.
The authors propose TS-SP (Two-Stage Speaker Preservation): first adapt the native audio encoder of Qwen2.5-Omni-7B with speaker identity supervision under LoRA, then freeze that encoder and train only the language model with LoRA for pairwise speaker comparison, keeping all pretrained base weights fixed; on Vox1-O this lowers equal error rate from 7.01% for the Paired Loss Adaptation Baseline to 4.37% and minDCF from 0.915 to 0.611, with Vox1-E and Vox1-H also improving from 7.64% to 4.41% and from 19.31% to 9.43%, equal error rate staying within 4.31–4.79% under unseen prompts, while cross-domain CN-Celeb equal error rate is comparable to the baseline (16.74% versus 16.49%) with lower accuracy at the native decision threshold.