Benchmarking open-source automated thigh muscle MRI segmentation algorithms
Synopsis
This preprint benchmarks eight open-source thigh muscle MRI segmentation tools on the MyoSegmenTUM, AIPS and Sheffield base datasets plus derived pathological and augmented out-of-distribution sets, combining Dice, Jaccard, Hausdorff, boundary IoU and inter-slice Dice ratio with qualitative usability review, and finds that domain-specific U-Net models, especially MuscleMap, generally outperformed foundation-model and newer general-purpose approaches, that adding SAM variants usually degraded rather than improved segmentation quality, and that most tools showed reduced accuracy on pathological cases.
Fig 1. Images from Sheffield dataset sample Aug-2 demonstrating poor labeling over the course of several muscles. The labeled muscles are shown in red. The left column shows the individual muscle label, while the right column shows all labels. The top row shows adductor brevis at slice 718, which is poorly labeled. The middle row shows gracilis at slice 718 with questionable labeling, while the bottom row shows gracilis at slice 738 with incomplete segmentation. The poor labeling of both muscles was present across multiple slices.
medRxiv · Page 5Interpretation
The authors conduct an independent head-to-head evaluation of eight open-source thigh muscle MRI segmentation tools on three base datasets plus a derived pathological subset of 12 bilateral thigh MRI volumes from four subjects with neuromuscular disease and an augmented out-of-distribution set; six tools run independently while two SAM variants require point or bounding-box prompts. The paper states that no independent head-to-head comparison for thigh muscle MRI had appeared in the literature, and that existing surveys either predate key foundation models or compare methods only qualitatively without quantitative benchmarking. The benchmark spans multiple datasets and several continuous metrics (Dice, Jaccard, Hausdorff, boundary IoU, inter-slice Dice ratio) alongside qualitative image review; Table 4 reports best Dice per algorithm per dataset with ordinal ranks.
Domain-specific U-Net models, especially MuscleMap, generally outperformed foundation-model-based and newer general-purpose architectures, and the thigh-specific MuscleMap model unexpectedly underperformed its whole-body counterpart. This provides a first shared-benchmark quantitative ordering for thigh muscle MRI tools, suggesting broader training data can compensate for reduced task specificity. Table 4 reports, for example, MuscleMap whole body with Dice 0.861 on AIPS, 0.825 on MyoSegmenTUM, 0.796 on augmented, 0.749 on pathological and 0.697 on Sheffield, versus Dafne at 0.072, 0.068, 0.020, 0.071 and 0.017 respectively.
Adding SAM variants to existing tools most often degraded segmentation quality, with a systematic pattern of including the femur or bone in muscle labels even when prompts came from well-performing tools. The work quantifies the effect of SAM augmentation across configurations and links it to a concrete qualitative failure mode. Both qualitative review across MyoSegmenTUM and quantitative SAM-variant experiments on MuscleMap WB and Dafne show a general trend of degradation with occasional improvement in poor performers; the authors report that SAM variants did not enhance any of the best-performing algorithms.
Most algorithms performed worst on the pathological dataset, and some degraded on augmented out-of-distribution data; the pattern was not uniform, with MuscleMap and MuSeg showing little to no Dice drop while MedCLIP-SAMv2 lost roughly half its Dice under augmentation alone. This divergence isolates positioning and orientation, rather than unfamiliar anatomy or disease presentation, as a variable driving the augmentation drop, and shows aggregate metrics can obscure poor performance on the muscles most affected by disease. Table 4 ranks datasets per algorithm by best Dice; the pathological subset contains 12 volumes from four patients, and the augmented set is derived from four diseased and six healthy subjects with hand-checked, anatomically plausible images.
Perspective
The immediate use of this benchmark is to help clinical researchers, neuromuscular imaging scientists and translational teams choose and operate open-source tools for volumetric analysis, radiomics and fat fraction quantification in thigh muscle MRI; the authors recommend pairing any automated segmentation with visual inspection and, where boundary precision is clinically critical such as radiotherapy planning, manual review and correction. The evaluation framework, augmentation pipeline and datasets are publicly available to support periodic re-evaluation as the field evolves.
Ground truth in the three base datasets was produced by single expert annotators (with radiologist supervision where reported), with no reported inter- or intra-rater reliability; some Sheffield labels are noted as incomplete or questionable, so accuracy values should be read as deviation from that reference. Not every muscle in every dataset was evaluated and results were not stratified by sex; the data do not include Africa or the Americas. Tools expect different input modalities (fat fraction, water, Dixon), so direct algorithmic comparison has inherent asymmetry. Table 3 marks Multimodal-Multiethnic as trained on MyoSegmenTUM, so its results on MyoSegmenTUM and derived sets can reflect data leakage. Dafne was run via script without its interactive correction workflow, so its numbers reflect uncorrected output. Figures 1 to 11 are removed from the loaded markdown with captions only, so readers checking bone inclusion or label discontinuity should consult the original figures. The rapidly evolving tool landscape makes this a snapshot, and whether the trend toward worse pathological performance generalizes to other datasets remains an open question.
