Skip to main content
Back to timeline
arXivSource publication:

Cross-meeting speaker attribution: ThyVoice leads all evaluated commercial cascades at a 47.13 mean SI-cpWER, while per-meeting leader ElevenLabs gains about 30 percentage points of attribution error on CHiME-6

Synopsis

The work introduces SI-cpWER, a metric that scores cpWER under one corpus-global speaker-ID map, and evaluates five commercial diarize-then-identify cascades, two open baselines, and the end-to-end reference system ThyVoice on CHiME-8 NOTSOFAR (clean and noise-augmented) plus CHiME-6, finding that requiring one persistent identity across meetings changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions, with a full-panel mean of 47.13 versus 54.75 for the next system.

AI-generated editorial illustration: Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

Interpretation

The paper proposes SI-cpWER, cpWER computed under one global speaker-ID map for the whole corpus, to measure whether the same person retains one identity across meetings. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently within each recording, so they cannot measure cross-meeting identity consistency; SI-cpWER folds persistent identity directly into scoring. Scored on all 129 NOTSOFAR evaluation meetings (12 disjoint test speakers, 13.34 hours, 188,036 scored reference words, roughly 30% overlapped speech) and two CHiME-6 far-field sessions (56,886 reference words), using a Hungarian solver for the global one-to-one assignment while keeping sample boundaries intact.

Under persistent attribution the commercial ranking changes: ThyVoice has lower SI-cpWER than every evaluated commercial cascade in clean, noisy, and CHiME-6 conditions, with a full-panel mean of 47.13 versus 54.75 for the next system. On per-recording cpWER, ElevenLabs is the best commercial system in every condition, but that lead does not carry through to cross-meeting attribution; SE-DiCoW leads the full panel on both NOTSOFAR conditions while ThyVoice leads on CHiME-6. The main table reports condition-level values and means for every system; on CHiME-6 ElevenLabs' SI-cpWER is approximately 30.0 percentage points above its local cpWER, compared with 1.5 points for ThyVoice.

Layered diagnostics separate the upstream error surfaces of the final attributed record: lexical recognition, local speaker activity, per-recording attribution, and speaker clustering each constrain the outcome. Rather than a single end-to-end number, the paper reports complementary WER, DER/JER, IAR, identity-count, and clustering purity/coverage/overlap-recall evidence, showing how the evidence surface available to the identity layer shapes persistent attribution. ThyVoice's DiariZen md-v2 component has the lowest DER/JER in every condition (clean 18.32/23.65), with 96.2% cluster purity, 95.2% coverage, and 78.5% overlap recall, whereas one-speaker-per-word surfaces show 0.0% overlap recall and cannot remove hidden overlap from timing alone.

An overlap-policy ablation shows that using overlap audio for transcription and for identity evidence are separate decisions: on the 20-meeting clean NOTSOFAR subset, separation plus stream rematching achieves the lowest cpWER of 34.39 and SI-cpWER of 32.73, while using mixture-containing audio for identity matching raises SI-cpWER to 67.76. The subset comparison separates whether overlap may inform identity enrollment from whether it may inform recognition, pointing to identity-evidence gating rather than recognition quality alone. Four overlap policies are compared on the same 20-meeting subset with pooled enrollment throughout; the authors state this subset provides directional evidence and does not establish the size of the effect over the full corpus.

Perspective

The evaluation targets speech systems that turn meeting audio into durable, queryable records, in an English, single-channel setting without a pre-enrolled speaker roster: NOTSOFAR supplies natural office meetings with participant aliases that persist across meetings, and CHiME-6 supplies real far-field dinner-party sessions. SI-cpWER and the layered diagnostics can be used directly to compare diarize-then-identify cascades and end-to-end systems on identity persistence, and to test design choices such as overlap repair and enrollment-evidence gating. The authors note that future work should evaluate recurring speakers over real multi-day use and measure whether attribution errors change retrieved facts, summaries, decisions, or action-item ownership.

Persistence is simulated by processing recordings serially in a fixed order against one shared voiceprint store rather than observing the same people recur across days, so the results characterize identity consistency under that fixed sequence; longer timescales and robustness to recording order remain open questions. The diagnostics each score the component or provider output surface available for that system rather than one uniform interface, and values are not added across layers. The noisy condition is deterministic augmentation rather than deployed capture, CHiME-6 contributes only two sessions, and multilingual and code-switching settings are not covered. ThyVoice's enrollment ablation and the provider pooling contrasts are single, unpaired runs, and repeated ElevenLabs draws show a 6.80 spread in noisy SI-cpWER, so those comparisons should be read as descriptive. In addition, the task-level cost of attribution errors, namely whether a fact or commitment is actually attached to the wrong person, is explicitly left as future work.

Sources