EchoChat recasts empathetic spoken dialogue as a three-stage perception–mental-state–response pipeline, lifting EchoEval perception from 6.96 to 8.23 and reasoning from 6.61 to 8.43
Related research and updatesSynopsis
The work introduces EchoChat, which reformulates empathetic spoken dialogue as a structured cognitive reasoning process spanning perception, mental-state reasoning, and response generation, supported by the EchoDialogue-400K dataset of 400K single-turn samples plus 4K multi-turn sessions, Acoustic-Anchored Attention, and a stage-aware RL objective with Step-Decomposed Credit Assignment, together with the 1K-instance expert-annotated EchoEval benchmark; experiments report gains over 12 compared models with perception 8.23, reasoning 8.43, and response 9.26, and an average Spearman correlation of 0.83 between LLM-based scores and human judgments.
Figure 1 : Reframing empathetic dialogue as structured cognitive reasoning. Existing methods simplify empathy to direct response generation, leading to cascaded reasoning errors across intermediate stages. EchoChat instead models empathy as a structured process from perception to response, supported by EchoDialogue-400K and EchoEval, enabling cognitively grounded and emotionally aligned spoken interaction.
arXivInterpretation
It reframes empathetic spoken dialogue from a direct input-to-response mapping into a structured cognitive reasoning process with perception, mental-state reasoning, and response generation, grounding the middle stage in Theory-of-Mind principles. Prior systems such as ParaS2S explored perception-to-speech-response alignment but still bypassed the intermediate reasoning about what the user needs and why a support strategy is appropriate; this work models that layer explicitly and optimizes it jointly. The three-stage framing runs through data construction, training objectives, and evaluation, with EchoEval reporting perception, reasoning, and response scores; the reasoning average of 8.43 exceeds the best compared model's 6.61.
It builds EchoDialogue-400K: 400K single-turn samples plus 4K multi-turn sessions of five turns each, about 9.8K hours of user-query audio and 3.3K hours of response audio, averaging 24 query words, 278 reasoning words, and 79 response words per sample, with prosody labels, intensity, sarcasm, and mental-state annotations. The paper's comparison table shows EChat-200K, EmotionalQA, ParaS2S, and OpenS2S all use coarse emotion labels and mostly lack open release, reasoning, or mental-state annotation; this dataset combines multi-turn, reasoning, mental state, prosody, intensity, and sarcasm coverage. Scale, audio hours, and annotation dimensions are reported in the main text and appendix statistics, and the pipeline includes multi-LLM multi-prompt generation, TTS synthesis, and consistency filtering with emotion2vec and a WavLM-based classifier.
It proposes Acoustic-Anchored Attention and Step-Decomposed Credit Assignment: the former maximizes attention mass from perception-stage tokens to acoustic tokens via an auxiliary SFT objective, and the latter decomposes stage-level rewards into token-level advantages across four spans (perception, reasoning, strategy, response) instead of GRPO's single broadcast scalar. Standard GRPO rewards only at the trajectory level, making intermediate errors hard to localize; cascaded gating requires earlier stages to be reliable before later stages are rewarded, and SDCA further assigns credit to the corresponding reasoning spans. Ablations show AAA raises emotion accuracy from 65% to 77%, intensity accuracy from 72% to 91%, and prosody score from 3.27 to 3.96, with a 136% increase in attention mass to acoustic inputs during perception; in the error-propagation table EchoChat reaches conditional correctness of 0.71 and 0.68 and an error ratio of 0.21, better than baselines.
It releases EchoEval: a 1K-instance benchmark of human-written queries recorded by 20 professional actors, covering implicit, audio-dominant, high-arousal, modality-conflict, and emotional-shift scenarios, with perception and ToM reference annotations from psychology-trained annotators. Existing benchmarks focus mainly on response-level quality and lack staged assessment of the full perception–reasoning–response pipeline; EchoEval provides per-stage 0–10 and 1–5 scoring protocols with hard-cap rules. Benchmark statistics report 1,000 utterances, 220 each across five scenarios (120 for emotional shift), averaging 29 words and 13 seconds of audio; LLM-based scores show an average Spearman correlation of 0.83 with human judgments.
Perspective
The framework targets integrated perception–reasoning–response modeling for empathetic spoken dialogue, suited to settings that require inferring a user's latent mental state from prosodic and semantic cues and generating emotionally consistent speech responses, such as the implicit, audio-dominant, high-arousal, sarcasm and other modality-conflict, and within-utterance emotion-shift cases covered by EchoEval. The public release of EchoDialogue-400K and EchoEval offers reusable resources for training and staged evaluation, and AAA and SDCA, as training-objective mechanisms, can inform other multi-stage speech reasoning systems. The paper explicitly places persistent user memory in long-term conversations, real-time interaction, and large-scale personalization outside the scope of the current study.
Evaluation centers on the self-built EchoEval, where perception and reasoning scores come from LLM judges; the paper supports their agreement with human judgments via an average Spearman correlation of 0.83, while external validation is limited to competitive results on VStyle. AAA's mechanism evidence is presented as a 136% increase in attention mass, and SDCA's benefit appears in ablations as small differences such as perception 8.20 versus 8.23 and reasoning 8.44 versus 8.43, so stability across base models and languages remains to be seen. Queries are generated by multiple LLMs and audio synthesized by TTS, whereas EchoEval is human-written and recorded by professional actors; how this distribution difference affects extrapolation is worth watching. Long-term memory, real-time interaction, and personalization are listed as out of scope, and their effects remain unclear.
