Arti-JEPA: Adapting a Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
Synopsis
The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.
Figure 1 : Arti-JEPA framework overview. Left: domain-adaptive pretraining on the largest available pool of vocal-tract real-time MRI, approximately 50 hours of high-frame-rate video, using a masked latent-prediction objective. Right: the resulting encoder is frozen and evaluated on a set of downstream tasks that probe how far the learned representation generalizes.
arXivInterpretation
It introduces a label-free recipe, Arti-JEPA, that continues V-JEPA 2's masked latent-prediction objective on about 62 hours of unlabelled rtMRI vocal-tract video from 92 speakers, with an fps-agnostic resampling path that places heterogeneous 61-99 fps sources on a 50 fps grid, and with label-free collapse diagnostics monitored throughout. Neither JEPA nor VideoMAE video models had previously been adapted to articulatory rtMRI; this work re-points a generic video prior onto the articulatory manifold at low cost (about 10 GPU-days) rather than training from scratch. The pretraining pool combines about 19.3 hours from USC 75-Speaker with about 42.6 hours of an in-house longitudinal corpus, totalling about 61.9 hours, 9,353 videos and 92 speakers; collapse monitors settle at feature std about 1.42, effective rank about 40/1024 and mean absolute cosine about 0.44, with the loss on an about 0.40 plateau, which the authors read as no collapse.
It builds a phoneme recognition and classification benchmark on the frozen representation, showing that a temporal video prior decisively outperforms per-frame image encoders and that, at matched adaptation, latent prediction (V-JEPA) is at least as strong as pixel reconstruction (VideoMAE), with the edge on fine-grained phonemes. It poses the latent-versus-pixel question as a controlled comparison, benchmarking six frozen encoders on the same ViT-L backbone and the same probe. On frame-level cross-domain decoding, adapted Arti-JEPA reaches kappa 0.505 in-domain and 0.352 cross-domain, against 0.362 and 0.156 for public pretrained V-JEPA 2; on segment-level classification Arti-JEPA leads on vowels, consonants, manner and place, with cross-domain vowel macro-F1 of 0.516, while stock encoders lose roughly half of macro-F1 on the unseen speaker and the adapted ones about a tenth.
The benefit of domain adaptation is task-dependent: it roughly doubles cross-domain phoneme kappa but helps fluent-versus-disfluent binary classification only marginally, and classifying disfluency type is substantially harder than binary detection. It carries the same frozen representation to typical speech and to two clinical populations, directly comparing how much adaptation buys on different tasks rather than reporting a single task's gain. For binary stuttering detection Arti-JEPA T-SSL reaches 0.817 macro-F1 against 0.789 for the generic V-JEPA 2 it was adapted from, with all four encoders inside a 0.03 band; three-class disfluency typing drops to pooled macro-F1 0.381 and kappa 0.110 against a three-class chance level of 0.333, with recall collapsing to 0.183 for repetitions and 0.318 for prolongations.
The glossectomy transfer analysis shows that a marked performance drop is already present preoperatively, that postoperative effects vary substantially across speakers, and that coarse articulatory structure is more robust than fine phoneme identity. It applies frozen probes in an inference-only setting to paired pre- and postoperative acquisitions, decomposing the drop into speaker-OOD, scanner/protocol and pathology/surgery rungs, and reports which phonemic contrasts collapse or are reorganised. Vowel kappa falls from 0.543 in-domain to 0.470 for the typical out-of-domain speaker and to 0.282 preoperatively for the glossectomy participants, with surgery adding only about 0.029 pooled kappa; place of articulation is most robust, with preoperative kappa 0.545 exceeding the typical-OOD anchor of 0.408, and the robustness ordering is place > manner > consonant identity > vowel identity. Spk1 and spk2 improve postoperatively on vowel and consonant identity, while spk3, who had the most extensive resection, declines sharply on all four tasks.
Perspective
The work targets English-language recordings, single grayscale mid-sagittal slices at roughly 84x84 to 104x104 pixels, with the encoder used frozen alongside light probes, and it applies to phoneme read-out in typical speakers, fluency detection in stuttering, and pre/postoperative transfer analysis after glossectomy. The authors frame the value as a reusable measurement tool for articulatory and clinical speech science where labels remain scarce; directions they leave open include rerunning the adaptation at larger compute, adding matched in-domain controls for the glossectomy cohort, extending to child speech and languages beyond English, and comparing audio-only with rtMRI classifiers to establish what articulatory imaging adds beyond acoustics.
A careful reader would still watch several things: the glossectomy cohort has only three participants, only one with an extensive lingual resection, so no systematic mapping between resection extent or location and which contrasts become less recoverable can be established; pre- and postoperative sessions differ in stimulus composition, confounding per-speaker pre-to-post differences, which is why the authors rely on the segment-level task. In the stuttering part, memory constraints cap the probe input at 32 frames per event, so longer disfluencies may be under-sampled, and class priors computed over training speakers do not match a held-out speaker, which drives the low recall in type classification. The acoustically conditioned latent-rollout attempt is reported as a negative result: over a 0.64 s clip the next-latent objective can be solved by extrapolating video alone, so making the acoustics necessary rather than merely available remains an open problem. The text also refers to figures and appendix details by reference, so readers wanting the per-class confusion structure should consult the original figures.
