Only content delivered through the publication boundary on this date is included.
arXiv The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.
The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.
The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.
The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.