BioGait-VLM reaches 68.1% accuracy on an 8-class gait benchmark and is chosen as the better model in 69.2% of blinded expert cases
Synopsis
The work proposes BioGait-VLM, a tri-modal (RGB vision, language, explicit biomechanics) framework built on a frozen InternVL3.5-1B, combining a Temporal Evidence Distillation branch with a Biomechanical Tokenization branch that projects 3D skeleton kinematics into language-aligned semantic tokens; on a subject-disjoint 8-class benchmark of 1,181 clips formed by merging public GAVD with a newly collected 30-patient degenerative cervical myelopathy (DCM) cohort, it reaches 68.1% accuracy and 52.9% macro F1 (98.1% F1 on DCM, 80.4% on Abnormal gait), and in a blinded review of 52 held-out clips four DCM-domain experts selected it as superior in 69.2% of cases (36/52, p<0.001).
Fig. 1. Overview of BioGait-VLM. The framework integrates three streams: (Top) Temporal Evidence Distillation (TED) for rhythmic dynamics; (Middle) Visual Con- text; and (Bottom) Biomechanical Tokenization, projecting 3D kinematics into se- mantic tokens. Fused within a frozen LVLM, these modalities enable state-of-the-art classification and evidence-grounded clinical reporting.
· Page 4Interpretation
Introduces a Biomechanical Tokenization branch that extracts a 46-dimensional SKEL pose vector per frame with an off-the-shelf HSMR estimator, converts it via a template mapping into structured natural-language descriptions (e.g., "Frame t: [Pelvis] tilt=α° list=β° ... [R.Knee] flex=δ°"), and tokenizes them together with a clinical instruction prompt so the LLM can directly "read" joint kinematics. Prior skeleton graph networks such as ST-GCN output only class probabilities, and general LVLMs lack the ability to quantify fine-grained temporal abnormalities; this work injects kinematic values as semantic tokens into the pretrained semantic space, removing the need for a separate skeleton encoder. Ablation shows that removing the biomechanical branch degrades performance on structurally defined pathologies (Parkinson's, DCM); in the blinded expert study the full model scores 4.23 on evidence grounding versus 2.31 for the strongest baseline.
Adds a Temporal Evidence Distillation (TED) branch that uses 32 learnable motion queries and a 3-layer, 4-head Transformer decoder to aggregate frame-level visual features via cross-attention into a global dynamic descriptor, capturing rhythmic anomalies such as tremor or festination. Standard VLMs typically process video with naive mean-pooling, which can obscure the fine-grained rhythmic cues needed for gait diagnosis; TED uses queries as latent temporal anchors that aggregate frame features by relevance to motion patterns rather than visual appearance. In the ablation, removing the TED branch causes a sharp performance drop, particularly in the Myopathic class, indicating that static VLM features struggle to capture rhythmic anomalies such as waddling.
Collects and curates the DCM clinical gait dataset: 30 patients and 239 video recordings from outpatient neurosurgical visits at a tertiary center at Washington University School of Medicine, labeled by an attending spine surgeon's diagnosis, captured at 30 fps and 1080p with a standardized sagittal protocol along a 10-meter corridor while retaining realistic clinical background clutter. Public datasets such as GAVD often lack specific spinal pathologies; this cohort specifically targets DCM, a condition with subtle and often misdiagnosed gait anomalies (spasticity, imbalance, increased cadence), filling that pathology class. Collection followed HIPAA and IRB guidelines with a dual-stream de-identification pipeline (skeleton path retaining no identifiers, visual path with face blurring before feature extraction), and only de-identified derivatives were used for training.
Builds a leakage-aware, subject-disjoint evaluation protocol: GAVD uses 158 subjects (728 sequences) for training and 40 subjects (214 sequences) for testing, while DCM follows an 8:2 patient-level split (24 patients, 187 sequences training; 6 patients, 52 sequences testing), giving a combined 915 training and 266 testing sequences with zero subject overlap. Unlike the original GAVD partition, this protocol prevents the model from memorizing patient identities, making evaluation closer to real cross-patient clinical generalization. Under this strict protocol, SlowFast, TSN, and zero-shot Qwen3-VL-2B and InternVL3.5-1B all fall below 40% accuracy, while the full model reaches 68.1%.
Perspective
The result targets clinical gait assessment settings where standardized sagittal-view capture is available and training and testing are split by subject; the method uses a frozen InternVL3.5-1B backbone and trains only the temporal decoder and a linear head, keeping compute needs relatively modest. For clinicians, its value lies in producing both a classification and a textual rationale citing concrete joint angles, step times, and swing percentages, usable as an aid for screening and documentation; for researchers, it offers a workable paradigm for injecting explicit biomechanics into vision-language models, plus an 8-class benchmark including a DCM class and a leakage-aware split protocol.
The DCM cohort is small (30 patients, 52 held-out clips), and macro F1 may vary substantially on underrepresented classes, so the stability of the 68.1% accuracy and 52.9% macro F1 still needs testing on larger, multi-center cohorts. The tokenization encodes only kinematic parameters such as joint angles, not kinetic quantities like forces and moments, and the parameters come from monocular pose estimation rather than direct measurement, so they may be affected by single-view depth ambiguity and occlusions such as loose clothing. The authors have begun a Phase 2 collection using OpenCap, a two-smartphone multi-view markerless motion capture system, aiming to replace estimated parameters with verified musculoskeletal sequences. In addition, the blinded expert review is a focused qualitative pilot in which each of four evaluators reviewed 13 cases, so the generality of its conclusions awaits larger-scale assessment.
