DuplexJev replaces autoregressive decoding with single-token read-out, answering 80 decisions over eight utterances in about 0.1 s on an 8-GPU node while a frozen LLM also hears gender and emotion
Related research and updatesSynopsis
The authors propose DuplexJev: ASR-encoder hidden states pass through a small connector into a frozen LLM, and each closed question is read as a single-token distribution over its options with nothing decoded, letting an 8-GPU node answer 80 decisions about eight utterances in about 0.1 s; with a last-layer connector spoken QA stays close to reading the transcript (90% vs. 91%), while a cross-attention connector brings gender and emotion accuracy to 90% each (from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%), training decisions by cross-entropy on the read-out answer token and keeping distillation for content.
Figure 2: R3 from the R2 connector, same data and schedule. Left: gender on 800 real utterances under distillation (flat) and answer-token supervision, for A and B. Right: emotion under answer-token supervision, with ZJU-ML (dashed).
arXivInterpretation
DuplexJev recasts the many small, closed decisions of full-duplex voice agents as single-token read-out: ASR-encoder hidden states go through a small connector into a frozen LLM, and each question is read directly as a single-token distribution over its options, with nothing decoded. Relative to current systems that answer such decisions by slow autoregressive decoding, this replaces token-by-token generation with one read-out and keeps encoders and LLMs interchangeable. The abstract reports an 8-GPU node answering 80 decisions about eight utterances in about 0.1 s, and states that weights, training recipe, a batched-inference pipeline and a bilingual spoken-QA set are released.
The connector design determines whether paralinguistic information survives: a last-layer connector keeps spoken QA close to reading the transcript (90% vs. 91%), whereas a cross-attention connector reaches 90% on both gender and emotion (up from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%). This offers a selectable design point between transcript content and speaker attributes, rather than optimizing transcript QA alone. The abstract provides these paired accuracy comparisons but no dataset sizes, statistical tests, or per-decision breakdown.
For training, decisions use cross-entropy on the read-out answer token instead of the usual transcript distillation, with distillation kept for content. The authors note that the teacher in transcript distillation never hears the voice, so that supervision cannot carry paralinguistic cues, whereas single-token cross-entropy acts directly on the decision read-out. The abstract supports this design choice with the method description and the accuracy changes above, without reporting ablation details.
Perspective
The work targets small, closed decisions in full-duplex voice agents, such as choosing among fixed options rather than open-ended generation; its setting is an ASR encoder plus a frozen LLM plus a small connector, served in batches on an 8-GPU node. The abstract states that encoders and LLMs are interchangeable and releases weights, training recipe, a batched-inference pipeline and a bilingual spoken-QA set, so it can be used directly for reproduction, backbone substitution, and transfer to similar closed decisions. The applicable premise is that a decision can be expressed as a single-token read-out over options and that the latency goal is mainly batched throughput.
The abstract does not state the size, language coverage or annotation scheme of the spoken-QA set and decision tasks, nor the hardware and batching configuration behind the latency measurement; the gender and emotion gains come from the cross-attention connector, and its difference from the last-layer connector on spoken QA (83% vs. 90%) suggests a trade-off between content and paralinguistic attributes whose curve is not yet clear. The abstract also reports no statistical significance, error-type distribution, or cross-encoder and cross-LLM reproduction results, which are open questions a reader would want to confirm before adoption.
