Duplex-MPE tests 2,000 multi-party dialogue scenarios and finds that talking more is not answering better, with MiniCPM-o 4.5 leading three of four capabilities
Synopsis
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
Interpretation
The benchmark turns the decision of whether to speak into an observable speech-control test: each scenario contains exactly one unresolved direct request T plus silence-requiring turns N1 (Aria mentioned but not asked), N2 (addressed to another human or device) and N3 (speech with no designated addressee), and an N4 event in which a human first asks Aria a question (N4Q) and a later utterance (N4R) answers it or states that no answer is needed, testing whether the assistant stops speaking. Existing benchmarks largely centre on a designated user and evaluate turn-taking, interruption handling or multi-round dialogue, or treat non-addressed speech as overlap to be filtered; Duplex-MPE makes every participant contribute to the dialogue state, treats the assistant as one of several possible addressees, and observes the waveform the model chooses to emit. The paper fixes six numerators and denominators in a table before examining results, scores four capabilities on separate denominators without forming an aggregate, generates scripts with Claude Opus 5 from five attributes (setting, activity, relationship, device context, register), synthesises turns with Qwen3-TTS, and validates sampled data and outputs with six reviewers.
The paired explicit-versus-implicit design shows that a transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often when the assistant is named than when it is not, while the five speech systems show no statistically significant paired response-rate difference. Prior work usually treats addressee recognition as a classification task over a supplied utterance; here it becomes a latent decision controlling whether an end-to-end model speaks, with the scenario task and gold answer held fixed while T's addressing form is reversed. The paper applies the exact McNemar test, and all five speech-system p-values exceed the 0.05 threshold (the smallest is reported as above it); the authors note the pattern could reflect either successful inference of the addressee from context or a failure to recognise the name as an addressing cue.
Separate scoring exposes distinct behavioural profiles: under explicit addressing Freeze-Omni has the highest response presence but the lowest fresh-onset response rate, FLM-Audio has response presence similar to MiniCPM-o yet a far lower conditional answer accuracy, and Freeze-Omni's near-perfect response presence coexists with silence preservation in only a small share of silence-requiring windows. A single response rate conceals what produced the behaviour; the paper separates a new decision to speak from speech already underway via fresh-onset response rate versus response presence, and separates selective responding from indiscriminate speech via silence preservation. The paper reports that under explicit addressing Freeze-Omni has the highest response presence and the lowest fresh-onset response rate, that FLM-Audio produces only a small number of correct answers among response-present requests, and that Freeze-Omni preserves silence in only a small fraction of silence-requiring windows; the silence-preservation ordering is unchanged across 0, 100, 300 and 500 ms duration thresholds.
Floor release is a coverage-conditioned outcome: under explicit addressing MiniCPM-o produces speech in most N4 windows, in some of which it stays silent during the question and starts speaking within three seconds after it ends, and only events still speaking at N4R onset enter the answering-window yield denominator; Freeze-Omni has speech in all windows but most are continuations from before the question, leaving only one event in that denominator. Stopping is separated from overall N4 performance, and a brief pause is explicitly insufficient—the model must remain silent from the deadline through the end of the observation window. The paper reports the explicit-condition answering-window yield denominators for each model, notes that FLM-Audio is still speaking three seconds after N4R begins in some scored events and that in some of Moshi's failed events the model is silent at the deadline but speaks again before the observation window ends, and does not display a rate for Freeze-Omni because only one event qualifies.
Perspective
The benchmark targets full-duplex speech systems: a model must receive continuous room audio, decide when to speak and keep listening while speaking, so non-full-duplex models are not scored and instruction sensitivity under a segmented interface is reported only as a non-scored probe. It is meant for settings such as meetings, living rooms and cars where the assistant Aria is one of several possible addressees and other participants' speech can both supply context for a later answer and resolve a request already addressed to the assistant. The authors state that the same taxonomy and automated generation and evaluation pipeline support larger or fresh evaluation draws, and they plan to release data and code before December, including scenario manifests, per-turn audio and construction boundaries, the duty preamble, scheduler and detector settings, per-scenario seeds and scoring code.
Many numbers in the loaded text render as blanks, including the six quantities in Table 3, most cells of Tables 9 through 21, and the dataset duration statistics, so this summary relies on the qualitative statements in the prose and the few figures stated explicitly (such as the 64.3 percentage-point difference for Gemini 3.1 Pro, the smallest speech-system McNemar p-value exceeding the 0.05 threshold, and the unchanged ordering across four duration thresholds). A careful reader would still watch several things: the explicit and implicit versions differ not only in T's wording but also in context edits and N4 realisations, so the authors compare response presence between complete paired scenarios without attributing the difference solely to the presence of the name; the speech systems' response rates change little between addressing conditions, a pattern that could reflect either successful addressee inference or a failure to treat the name as an addressing cue, which this study does not separate; the scenarios are synthetic English conversations that may inherit biases from the generating and speech-synthesis models; and Freeze-Omni's yield rate rests on a single eligible event, so its conditional rate should not be extrapolated.
