Public articles linked to the same research event.
arXiv This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.
This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.
This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.
This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.