Skip to main content

Daily report

AI and science frontiers · 2026-09-09

Only content delivered through the publication boundary on this date is included.

arXiv

Arti-JEPA: Adapting a Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.