NTU team finds speaker-verification EER does not track human voice similarity, and an embedding-dimensionality bottleneck lifts alignment correlation from 0.08 to 0.74
Synopsis
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Interpretation
Perceptual alignment is governed mainly by the training objective rather than verification accuracy: within every model condition the best and worst objectives differ by more than a factor of three in human perceptual alignment without an EER cost; prototypical metric losses are highest (0.395-0.499), classification losses lowest (0.080-0.298), and AM-Softmax is last in every condition. The community had assumed the lowest-EER speaker verifier is the best proxy for human listening; this work tests that premise against a common human-rating benchmark and provides a counterexample: on ECAPA, swapping AAM-Softmax for Angular Prototypical raises alignment from 0.111 to 0.406 at a tied-best EER of 1.47%. 35 training conditions (five model conditions x seven objectives x three seeds), with Spearman rank correlation between embedding cosine similarity and mean listener ratings computed over 16,905 different-speaker VoxSim pairs; EER is reported on VoxCeleb1-O using raw cosine similarity.
Embedding effective dimensionality (d_eff) is a geometric indicator of perceptual alignment: within the five model conditions the Spearman correlation between d_eff and alignment ranges from -0.85 to -0.99, reaching -0.95 when pooled across the 35 conditions, with permutation tests rejecting zero correlation in four of five conditions (ReDimNet-B2 excepted). Prior work documented that objectives reshape representation dimensionality but did not connect it to human perceptual alignment; this study links the existing finding that human voice similarity judgments can be captured by low-dimensional spaces to embedding geometry, and notes d_eff is computed from speaker centroids alone with no human ratings. d_eff is measured by the participation ratio of the PCA eigenvalue spectrum of speaker centroids, after L2 normalization and averaging by speaker to suppress utterance noise; ECAPA-TDNN's 21 runs fall tightly along the diagonal and the seven condition means correlate at -0.99.
Explicitly compressing the embedding dimension improves perceptual alignment: on ECAPA-TDNN, narrowing the dimension from 192 to 3 raises alignment for all four objectives, with AM-Softmax going from 0.080 to 0.738, making it the best-aligned model in the study. This turns d_eff from an observational correlate into an actionable training lever, showing alignment can be manipulated geometrically rather than only measured after the fact. Each dimension setting is trained with four objectives and three seeds; widening from 192 to 512 changes EER, d_eff and alignment almost not at all, showing the nominal dimension is only an upper bound; compressing to 10 and 3 raises EER to 13-37%, and AAM-Softmax fails to converge in two of three seeds at dimension 3.
The decoupling of EER from human judgment also holds for public state-of-the-art checkpoints: under the ECAPA condition four public checkpoints land in the low-EER but low-alignment region, contradicting the community expectation that a better verifier is a better proxy for human perception. This suggests the widespread practice in TTS and voice conversion of using speaker-embedding cosine similarity as an automatic timbre-similarity metric needs its selection rationale revisited. Comparison covers public checkpoints and author-trained models on the same VoxSim pairs; across the 35 conditions the overall EER-alignment correlation is -0.13, the sign is unstable within single model conditions (-0.14 to 0.43), and no permutation test rejects zero correlation.
Perspective
The result targets speech generation and conversion pipelines that assess synthesized timbre similarity via speaker-embedding cosine similarity, and applies to English VoxCeleb1 speakers, VoxSim human ratings, and the five model conditions and seven training objectives examined; the authors release code, configurations and seeds so the setting can be reproduced and extended. Because d_eff needs only speaker centroids, it can serve as a geometric monitor during training or model selection, screening candidate embedding models when no human ratings are available.
The dimensionality bottleneck raises alignment while pushing EER to 13-37%, and AAM-Softmax fails to converge in two of three seeds at dimension 3, so how to set the trade-off between alignment and speaker discrimination in a real system remains open. The d_eff-alignment correlation does not pass the permutation test on ReDimNet-B2, so the robustness of this geometric indicator across architectures needs more conditions. In addition, human ratings come from 13 listeners at about two ratings per pair, and the text does not quantify how rating noise caps the correlation; VoxSim is limited to English VoxCeleb1 speakers, so cross-language and cross-dataset generalization remains to be observed.
