Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Benchmarking seven video foundation models on 32,847 videos from 1,888 participants, VideoPrism ranks first on 10 of 16 tasks while V-JEPA2-SSv2 reaches 85.3% AUC on flip palm

This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.