A frozen Sapiens pose foundation model detects climbing hold usage without climbing-specific training, reaching 90.2% event F1
Synopsis
The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
Fig. 1: Training-free hold-usage pipeline. Snowflakes ( ) mark the frozen, off-the-shelf foundation models, which receive no climbing-specific training. A promptable segmenter [ 6 ] returns the climber box, Sapiens [ 5 ] estimates top-down pose with 308 whole-body keypoints of which we keep the fingertips and toes, and the proximity of each tip to the given hold box ( eq. 1 ), after per-limb mutual exclusion and a short temporal pass, yields the usage events.
arXivInterpretation
Using fingertip and toe keypoints from a frozen Sapiens pose foundation model, together with a per-frame proximity test, per-limb mutual exclusion, and a short temporal-persistence rule, hold-usage events can be detected without any climbing-specific training. Existing approaches either train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle, offset from the fingertips and toes that actually contact the holds, and whose hands are occluded in roughly half of all frames; this work instead uses the foundation model's fingertip and toe keypoints while keeping the model frozen. Evaluated on the The Way Up dataset (22 videos, 10 athletes, two routes), with an event F1 of 90.2% on a held-out split, 89.8% under leave-one-participant-out cross-validation, and 79.9% over all 22 videos at any temporal overlap.
Under an identical protocol, the method exceeds reproductions of the YOLOv8-pose and ViTPose pipelines at every temporal threshold, with the margin widening under strict timing. It provides a direct comparison against existing pose-estimation pipelines under a unified protocol rather than reporting only its own absolute metrics. The comparison is based on the authors' reproductions of the YOLOv8-pose and ViTPose pipelines, reported at every temporal threshold.
An ablation shows that two intuitively helpful additions, dense foundation-feature change gating and body-part segmentation, both hurt performance, supporting a minimal, keypoint-only design. Contrary to the common practice of stacking more features or segmentation information, the result indicates that a simpler design is the right one for this task. The ablation compares performance with and without these two additions.
Standard coaching statistics computed from the automatic predictions track ground truth closely, turning ordinary single-camera video into reliable performance metrics with no instrumentation. It extends detection results to practically usable coaching metrics, not just event detection itself. Pearson r=1.00 for climb time and 0.94 for pace.
Perspective
The result targets climbing settings that use single-camera video and annotated holds, applicable to hold-usage detection for automated scoring, movement analysis, and assistive systems; the method is validated on two routes and 10 athletes in the The Way Up dataset, with generalization reported through a held-out split and leave-one-participant-out cross-validation. For readers who want to avoid climbing-specific training and directly reuse an off-the-shelf pose foundation model, the work offers an actionable minimal design path and shows that automatic predictions are sufficient to support standard coaching statistics such as climb time and pace.
The work is validated on two routes and 10 athletes in the The Way Up dataset, and generalization to more routes, climbing gyms, and filming conditions remains an open question; hands are occluded in roughly half of all frames, so the effect of occlusion on detection and how the temporal-persistence rule behaves under different occlusion patterns deserve further observation; the mechanism behind why dense foundation-feature change gating and body-part segmentation hurt in the ablation is not developed in the text; and hold-usage detection depends on annotated holds, so how hold annotations are obtained and at what cost for practical deployment remains an open question.
