Public articles linked to the same research event.
arXiv The authors propose DuplexJev: ASR-encoder hidden states pass through a small connector into a frozen LLM, and each closed question is read as a single-token distribution over its options with nothing decoded, letting an 8-GPU node answer 80 decisions about eight utterances in about 0.1 s; with a last-layer connector spoken QA stays close to reading the transcript (90% vs. 91%), while a cross-attention connector brings gender and emotion accuracy to 90% each (from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%), training decisions by cross-entropy on the read-out answer token and keeping distillation for content.
The authors propose DuplexJev: ASR-encoder hidden states pass through a small connector into a frozen LLM, and each closed question is read as a single-token distribution over its options with nothing decoded, letting an 8-GPU node answer 80 decisions about eight utterances in about 0.1 s; with a last-layer connector spoken QA stays close to reading the transcript (90% vs. 91%), while a cross-attention connector brings gender and emotion accuracy to 90% each (from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%), training decisions by cross-entropy on the read-out answer token and keeping distillation for content.
The authors propose DuplexJev: ASR-encoder hidden states pass through a small connector into a frozen LLM, and each closed question is read as a single-token distribution over its options with nothing decoded, letting an 8-GPU node answer 80 decisions about eight utterances in about 0.1 s; with a last-layer connector spoken QA stays close to reading the transcript (90% vs. 91%), while a cross-attention connector brings gender and emotion accuracy to 90% each (from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%), training decisions by cross-entropy on the read-out answer token and keeping distillation for content.
The authors propose DuplexJev: ASR-encoder hidden states pass through a small connector into a frozen LLM, and each closed question is read as a single-token distribution over its options with nothing decoded, letting an 8-GPU node answer 80 decisions about eight utterances in about 0.1 s; with a last-layer connector spoken QA stays close to reading the transcript (90% vs. 91%), while a cross-attention connector brings gender and emotion accuracy to 90% each (from 55% and 28%) with spoken QA dropping only 1 point (83% to 82%), training decisions by cross-entropy on the read-out answer token and keeping distillation for content.