Skip to main content
Back to timeline
arXivSource publication:

EvolvingAvatar lets a 3D head learn while it generates, cutting expression-statistics mismatch by up to 11.1% on OOD-Hard

Synopsis

EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.

AI-generated editorial illustration: EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

Interpretation

It introduces EvolvingAvatar, a causal generator that performs test-time training during interaction without target motion labels at deployment and without precomputed user FLAME trajectories. Existing generators use incoming observations as context but keep parameters fixed, which the authors identify as leaving conversational patterns unused as a learning signal; this work treats the dialogue itself as a self-supervised adaptation signal, adjusting the context-to-motion mapping during generation. The paper provides a task formulation, a region-structured causal FLAME codec, and an adaptive generator, with architecture, training objectives, and streaming inference documented in Appendix E; in ablations, No write raises ID expression P-FD from 16.79 to 21.09, supporting a contribution from online updates.

Dyadic Context Prediction builds a self-supervised inner objective from user video and both participants' audio; persistent fast weights accumulate updates within a conversation to guide motion generation, transient jaw adaptation responds to current articulation, and predicted speech activity controls how persistent adaptation guides motion. Compared with conditioning on avatar audio, dyadic audio, or dyadic audio with estimated user motion, this work splits the adaptive state into a path retained across intervals and a path discarded after one interval, and gates the adaptation term by activity. Matched-checkpoint comparisons show full TTT lowers ID expression P-FD by 20.4% against No write, 4.2% against Freeze-1, and 8.6% against No carry; neck P-FD instead rises 8.9% against No write, which the authors attribute to DCP not directly constraining each region's motion statistics.

It releases InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos, with aligned multimodal annotations and evaluation of speaking and listening across distributions. dialog3d-factory unifies both recording formats onto a shared timeline, preserves overlap and silence, and combines parameter-space and mesh-space metrics with human ratings. Data come from Seamless Interaction and RealTalk recordings, split into Train, Dev, ID, OOD, and OOD-Hard, totaling 455.95 hours of interaction and 4,365 source participant IDs; the authors note these metadata IDs do not establish globally unique people.

Experiments show improved conversational motion statistics and improvement as conversations unfold on the hardest out-of-distribution split, with human judgments favoring the method under distribution shift. The authors separate agreement with one recorded response from perceived conversational quality: DualTalk has lower MSE and LVE/MHD but lower motion coverage (SID), while overall-realism preference against DualTalk on OOD-Hard is 86%. On OOD-Hard, speaking/listening FDD is 19.09/18.13 versus 20.91/19.84 for ARTalk; across five equal-duration intervals expression P-FD falls by 2.7% on OOD and 7.1% on OOD-Hard while all four baselines worsen on OOD; preference over UniLS rises from 58% on ID to 66% on OOD and 72% on OOD-Hard.

Perspective

The result targets interactive 3D head generation: inputs are causally arrived user face frames and both participants' audio, outputs are 106-D FLAME motion (expression, neck pose, jaw pose), time is partitioned into 200 ms intervals, streaming emits five frames after the fifth arrival with a 160 ms buffering delay, and all state resets between conversations. It lets researchers compare speaking and listening on a shared timeline, evaluate motion statistics and perceptual quality across distributions, and make learning during a conversation a measurable object. The authors point to potential applications in virtual tutors and in avatars that offer mental health support alongside human clinicians, while stating that recorded conversations do not capture how users respond to generated motion; continual learning in live interactions, distinguishing stable expressive habits from temporary emotional responses, and clinician-guided assessment of perceived support are explicitly left as future work.

Worth watching: the paper itself notes paired error metrics measure agreement with one recorded response, whereas an interaction can admit several appropriate responses, so advantages on MSE, LVE, and MHD are not equivalent to perceived quality; UniLS retains lower FDD on ID/OOD, showing that matching coefficient statistics does not ensure equally accurate geometric variation after decoding. Neck P-FD rises 8.9% against No write, suggesting DCP adaptation does not directly constrain each region's motion statistics. Jaw flow matching, activity conditioning, and RGB comparisons are labeled exploratory across revisions, and RGB improves expression and neck statistics on OOD-Hard but worsens expression statistics on ID/OOD. The UniLS preference trend is described as descriptive because all three uncertainty intervals include equal preference. In addition, this evidence bundle is full text with figures rendered as tables, so graphical content such as Figure 6 cannot be directly verified; confirming the exact shape of the interval curves and preference intervals would require the original figures.

Sources