Benchmarking seven video foundation models on 32,847 videos from 1,888 participants, VideoPrism ranks first on 10 of 16 tasks while V-JEPA2-SSv2 reaches 85.3% AUC on flip palm
Synopsis
This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.
Fig. 1. Overview of the benchmarking framework for PD screening using video foundation models (VFMs). (A) Video Dataset: Our study utilizes a large-scale dataset from 1, 888 participants performing 16 standardized clinical tasks. (B) Evaluation Pipeline: Raw video data is processed through a suite of state-of- the-art frozen VFMs to extract latent representations. These embeddings are evaluated using a task-specific linear classification head to differentiate between PD and non- PD participants. (C) Task-Model Saliency: Systematic evaluation to investigate architecture-specific strengths.
· Page 2Interpretation
The study establishes the largest webcam-recorded Parkinson's disease video dataset to date, comprising 32,847 videos from 1,888 participants (727 with PD) across 16 standardized clinical tasks inspired by the MDS-UPDRS, organized into four clinical domains: upper-limb motor kinematics, visual speech kinematics, facial expressivity, and oculomotor/cervical/cognitive control. Prior computer-vision-based PD screening work relied on handcrafted features mimicking clinical observations or task-specific models that could not generalize across tasks; this dataset and task framework provide a unified basis for cross-task and cross-architecture comparison. Data were compiled from eight independent IRB-approved clinical and non-clinical studies (2017–2025); PD status was determined by self-report (445) or clinical confirmation (282); the cohort is predominantly White (1,366), with 350 participants not disclosing race.
Under a unified frozen-backbone protocol with a linear classification head, the seven video foundation models achieved AUCs of 76.4–85.3% and accuracies of 71.5–80.6% across 16 tasks, with Flip Palm and Open Fist yielding the highest AUCs (85.3%±0.2% and 84.3%±0.1%, respectively) and Flip Palm reaching a peak specificity of 90.3%±0.5%. This is the first study to compare multiple VFM architectures under the same clinical task framework, showing that frozen pretrained embeddings can capture clinically meaningful signs without task-specific fine-tuning. Results are based on participant-level train/validation/test splits (60%/20%/20%), with means and 95% confidence intervals reported across 30 random seeds when feasible; hyperparameter search was conducted on Weights & Biases with equal compute time allocated across all tasks and VFMs.
Task saliency is highly model-dependent: VideoPrism ranked first on 10 of 16 tasks, with its superiority most pronounced in visual speech kinematics and facial expressivity; V-JEPA2 variants dominated upper-limb motor tasks (e.g., Flip Palm, Extend Arm, Nose Touch), with V-JEPA2-SSv2 consistently outperforming base V-JEPA2; and TimeSformer performed best on rhythmic tasks like Finger Tapping. The comparative effectiveness of different VFM architectures across specific PD screening clinical tasks was previously poorly understood; this study reveals correspondences between pretraining paradigms (semantic-visual distillation, latent-predictive world modeling, divided space-time attention) and specific clinical domains. Conclusions are based on cross-model comparison of validation-set AUCs across 16 tasks; V-JEPA2-SSv2's additional fine-tuning on the SSv2 dataset is associated with its advantage on upper-limb coordination tasks.
Multi-view training did not significantly outperform single-view (p>0.70 for both accuracy and AUC), and oversampling strategies slightly but non-significantly degraded mean accuracy (74.51% vs 74.97%) and AUC (79.83% vs 80.54%), suggesting that frozen VFM embeddings are already robust enough to handle existing class imbalances. This ablation provides empirical guidance on whether multi-view aggregation and class balancing are needed in remote video screening, suggesting that salient diagnostic features are often captured within short segments. The single-view configuration without oversampling achieved a mean accuracy of 74.97% (std: 3.12%) and mean AUC of 80.54% (std: 2.42%) across all 16 tasks; multi-view and oversampling comparisons reported statistical test results.
Perspective
This study establishes a baseline for video foundation model-based remote Parkinson's disease screening, applicable to settings where standardized tasks are performed using everyday devices (smartphones, webcams) at home or under clinical supervision. Its frozen-feature protocol makes results directly usable for comparing the inherent representational power of different VFM architectures, providing guidance for researchers selecting models suited to specific clinical domains (e.g., upper-limb motor kinematics or orofacial manifestations). Code and anonymized structured data are publicly available for reproduction and extension.
The frozen evaluation protocol does not explore performance ceilings achievable through task-specific fine-tuning or parameter-efficient methods such as LoRA; experiments are restricted to open-weight models deployable locally; some PD diagnosis labels come from self-report and may introduce label noise; the cohort is predominantly White, and generalizability to more diverse populations remains to be validated. Additionally, the gap between high AUC and low sensitivity suggests decision thresholds may be non-optimal, and clinical screening utility may require improved model calibration and integration of multiple tasks and modalities.
