Skip to main content
Back to timeline
arXivSource publication:

REALM adds a delay prior and stochastic expression refinement to turn listening from a frozen stare into blinks and smiles on a robot

Synopsis

REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.

AI-generated editorial illustration: REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

Interpretation

REALM models listening as a delayed reactive process: the Reactive Gated Speaker–Listener Fusion module uses a delay-centered attention prior, implemented as a shifted ALiBi bias, to favor speaker representations near a nominal response lag, and a learned gate to adaptively weight the aligned speaker context against listener history. Prior methods explore speaker-conditioned generation, autoregressive prediction, and dyadic representation learning, leaving the temporal relationship to be discovered from data; REALM introduces an explicit delay prior and separates when to align from how strongly it should influence generation. A delay-sensitivity analysis on ViCo shows an 8-frame shift (about 267 ms) achieves the lowest error, FD, and rPCC for both expression and head pose, better than 0, 4, and 12 frames; ablation shows shifted attention alone reduces pose rPCC from 0.026 to 0.018, while without gating expression FD rises to 0.68.

The framework uses a coarse-to-fine decomposition: a coarse decoder predicts a base motion trajectory, and a refinement module adds audio-conditioned stochastic residuals only in the non-rigid expression subspace, leaving the rigid pose parameters unchanged by construction. Deterministic reconstruction objectives tend to average multimodal local variation into a smooth trajectory, while injecting stochasticity across the whole motion space can perturb stable rigid motion; REALM restricts refinement to expressions as a complement to discrete motion representations and generative modeling. Ablation shows the refinement module alone improves expression FD to 0.62 and reduces expression error to 4.11, with the full model at 0.56 and 3.91; because residuals act only on the non-rigid subspace, pose kinematics are unaffected.

On the ViCo and L2L conversational benchmarks, REALM attains the lowest expression and pose point-wise errors among the evaluated methods, with the most pronounced gains in the temporal transition metric. The paper reports expression error dropping from 8.71 to 3.91 on ViCo and reaching 6.50 on L2L, versus 8.71 and 13.41 for ListenFormer, suggesting stochastic refinement mainly aids inter-frame facial dynamics rather than static trajectory tracking. Evaluation covers point-wise accuracy, distributional realism (FD, FID), speaker–listener correlation (rPCC), and motion variability across five baselines (RLHG, DSPN, L2L, ListenFormer, UniLS); the paper also notes RLHG retains a lower pose FD on ViCo (0.72 vs. 1.01) and ListenFormer has expression variance closer to the reference on ViCo (0.142 vs. 0.133).

Generated motion is retargeted to an Ameca humanoid robot through a robot-specific mapping and relative motion calibration, and a perceptual user study gives REALM the highest scores on naturalness, audio-visual synchrony, contextual appropriateness, and overall preference. Most generative avatars remain in virtual environments; REALM uses inverse kinematic mapping, Savitzky–Golay smoothing, and calibration relative to a mechanical neutral pose to convert 3DMM coefficients into actuator controls, supported by a structural blink analysis of the recovered micro-dynamics. The user study uses a 1–5 MOS scale with concealed method identities, randomized order, and embedded attention checks; the paper reports statistically significant pairwise differences between REALM and the strongest baseline under a Wilcoxon signed-rank test, with 1.50M parameters and 0.02G FLOPs per forward pass.

Perspective

The work targets audio-driven listener motion generation over offline conversational windows, suited to digital avatars and humanoid robot facial expression settings such as telepresence and embodied conversational agents; for readers who want to reproduce or extend it, code and training scripts are provided in the supplementary materials, and the appendix documents the two-stage curriculum loss, architectural hyperparameters, and evaluation metric definitions. The physical deployment relies on a robot-specific retargeting pipeline including inverse kinematic mapping, temporal smoothing, and calibration relative to a mechanical neutral pose, with actuator-range clipping and safety constraints, so results apply to platforms with comparable actuators and safety envelopes.

The paper frames REALM as window-conditioned generation and explicitly does not infer an end-to-end streaming guarantee from the attention mask, so extending to streaming human–robot interaction requires explicit control of temporal information access and end-to-end latency. The nominal delay is a shared alignment prior that may not capture variation across individuals, and the authors list context-dependent delay priors as future work. Physical deployment relies on a robot-specific retargeting pipeline, and extending across embodiments will require morphology-aware mappings and validation of actuator limits on each platform. In addition, ARIG and Wang et al. are not reproduced under this protocol because public implementations were not located, leaving compatible comparison as outstanding work; this summary is based on the full paper and the homepage evidence bundle and does not include every figure or table detail from the appendix.

Sources