DuplexAgent coordinates full-duplex voice with delegated agents through six editable modules, and its recursive loop lifts fixed-core collaboration from 54.6 to 85.3
Synopsis
The work presents DuplexAgent, which writes the collaboration between a full-duplex interaction model and a delegation pool of reasoning LLMs and coding agents as six editable modules (Policy, Intake, Hold, Worker, Deliver, Memory), plus Duplex-Harness-RSI, a closed loop in which a simulator generates timed conversations and attributes failures, a reasoning LLM acts as Exam Planner to choose the next exam, and a coding agent acts as Harness Editor to revise the modules named by the diagnosis, with paired development and protection gates deciding retention; on VoiceChat and Qwen-Realtime the retained harness improves spoken knowledge, executable tool use, and fixed-core collaboration.
Figure 1: Benchmark results for DuplexAgent (blue), compared with other collaboration systems (gray). Higher is better on every panel. Intelligence reports MMSU accuracy, agentic tool use reports BFCL parallel-multiple and FDB-v3 Pass@1, and duplex reports FDB-v1 interrupt TOR. DuplexAgent is stronger on all three capabilities. All collaboration systems use the same delegation LLM. See Tables 4 and 5 .
arXivInterpretation
DuplexAgent makes the full-duplex collaboration workflow explicit as six editable modules, each with a clear input, decision, and observable effect, so a failure can be attributed to a bounded part of the harness. Existing systems implement useful instances of handoff, acknowledgement, and delivery control, but those decisions are generally coupled in system-specific heuristics; this work separates routing, task interpretation, pending-work conversation, delegation contract, delivery timing, and shared context into Policy, Intake, Hold, Worker, Deliver, and Memory, supported by a task ledger holding task identity, version, status, and delivery history. The method section gives a module table and a Snake-to-Tetris session showing that a late Snake result is recorded but cannot make an obsolete task current again; the experiments run separate improvement campaigns on two interaction models.
Duplex-Harness-RSI reuses the delegation pool that serves users: the Exam Planner adjusts the next exam from attributed failures and the repair archive, the Harness Editor proposes a testable hypothesis on declared editable surfaces, and development and protection gates decide retention by paired wins and losses. Compared with one-shot pooled editing and with sequential editing that omits per-round diagnosis and the repair archive, the loop connects diagnosis, repair, and verification while keeping the simulator, scoring rules, and protection set fixed. On VoiceChat the initial harness and Qwen-Audio-Agent both score 54.6, one-shot editing reaches 57.4, sequential editing 61.9, and the ten-round checkpoint 85.3, improving on the initial harness by 30.7 points and on sequential editing by 23.4 points; on Qwen-Realtime the initial harness scores 47.1 and reaches 84.8 after eight retained versions, a gain of 37.7 points.
On external benchmarks, DuplexAgent separates the decision to delegate from the content of the handoff, so spoken knowledge and executable tool use rise together rather than trading off as with description-only handoffs. MoshiRAG, Gander, Realtime-Venus-Omni, and Qwen-Audio-Agent pass a task title or a natural-language restatement to the delegate and therefore stay at or below the matching interaction model on BFCL/FDB-v3, where the scored object is the benchmark function name with its required arguments; DuplexAgent has Worker receive the official schema and the arguments carried by the user turn, Memory retain identifiers and constraints that a title or short restatement drops, and Deliver return only the current result in a speaking window. On VoiceChat BFCL average rises from 53.5 to 75.8 and FDB-v3 from 60.2/26.2/16.0 to 97.4/57.0/47.0; on Qwen-Realtime BFCL average rises from 82.0 to 90.8, parallel-multiple from 66.0 to 78.0, FDB-v3 argument accuracy from 47.8 to 55.0 and Pass@1 from 39.0 to 49.0; spoken knowledge rises from 54.2 to 80.6 and from 86.3 to 92.1 on OBQA, and from 36.1 to 80.5 and from 64.8 to 82.7 on MMSU.
The retained repairs show concrete rules learned per module, and rejected broad repairs show the gates constraining the scope of change. The repair archive keeps both accepted and rejected attempts, letting the Exam Planner and Harness Editor avoid repeating unsuccessful changes and build on repairs that already helped. Recorded local exams show: stating direct-answer categories explicitly reduced extra calls by 50%; resolving late cancellation against live task state reduced ignored cancellations by 81%, with replacement failures falling by 65% and later by 83.3%; earlier progress reporting reduced missing progress updates by 75%; preserving requested units and precision reduced wrong answers by 66.7%; speaking only after the user is silent reduced overlaps and concealed failures by 100% each; assembling the complete turn reduced missing-context failures by 60%.
Perspective
The results apply to voice-agent deployments where a full-duplex interaction model serves as the conversational entry and reasoning and coding work is carried out asynchronously by a delegation pool, covering open-weight and hosted interaction models and one or several delegates. The harness uses adapter declarations such as native function calling, textual delegation markers, context injection, and the treatment of interim tool outputs to choose a compatible receipt and delivery path. The simulator can generate training or evaluation conversations for other interaction systems, and external benchmark scores and simulator collaboration scores answer different questions and are reported separately. The improvement loop rewrites module instructions and executable rules while interaction-model weights stay fixed.
The gains stop at the harness: a repair can require silence before speech or a complete question before delegation, while the model's perception of overlapped speech and its own judgement that a turn has ended remain as they were. When the gates accept or reject a candidate, this remainder is still in the trace, showing which failures a module edit can absorb and which remain in the interaction model's own perception and timing. On Qwen-Realtime, Candor pause TOR improves from 0.278 to 0.231 while synthetic pause TOR moves from 0.074 to 0.147, so better task performance coexists with a less favorable result on one pause condition. On VoiceChat, irrelevance falls from 95.0 to 90.0, so recovering required calls also admits some unnecessary ones; on the fixed core, extra delegation rises from 3.9% to 6.6% while ignored control stays at 5.3%. The per-module percentages come from the local exam used to test each repair, and fixed-core scores combine the repairs the gates retained. The Qwen-Realtime campaign shows a late plateau of rounds in which further proposals no longer passed the gates. Readers interested in whether the interaction model's own perception and timing can be trained remain with the direction the paper leaves to agentic reinforcement learning.
