IntraStyler uses a contrastively trained style encoder to synthesize T2 MRI styles without predefined sub-domains, raising Extra-VS Dice from 0.61 to 0.81 with zero failures on CrossMoDA
Synopsis
The work proposes IntraStyler, a 3D unpaired image translation method that requires no predefined sub-domains, using a contrastively trained style encoder to extract anatomy-disentangled style embeddings from each target-domain image and conditioning a QS-Attn synthesis network via dynamic instance normalization; on the CrossMoDA benchmark (226 labeled ceT1, 295 unlabeled T2, 96-image T2 test set) it automatically discovers 7 style references instead of the 3 used by sub-domain methods, and downstream nnU-Net segmentation achieves the best Dice and ASSD on Intra-VS, Extra-VS, and Cochlea, with Extra-VS Dice of 0.81 and zero failures.
Fig. 1. (a) Intra-domain variability: images from the same domain exhibit diverse ap- pearances across institutes and scanners. (b) Comparison of style generation strategies.
· Page 2Interpretation
Proposes the first 3D unpaired cross-modality MRI image translation method that needs no predefined sub-domains, treating each target-domain image itself as a style reference. Prior methods either train separate synthesis networks per source-target sub-domain pair or condition a unified network on one-hot sub-domain codes, both depending on the availability and accuracy of sub-domain labels; IntraStyler removes that premise. Supported by the method description and the ablation design in which NoVar, MultiNets, Dynamic, and IntraStyler differ only in how intra-domain variability is handled while the downstream segmentation model is fixed to nnU-Net.
Designs a contrastive learning objective to train a 3D style encoder that extracts per-image style embeddings disentangled from anatomical content. Query and positive are patches from different locations of the same image (same style, different anatomy), while negatives are intensity-perturbed versions of the positive (same anatomy, different style), making the encoder sensitive only to style; trained only on artificial perturbations, it still discriminates real MRI styles. K-means clustering with PCA visualization shows images within a cluster share consistent styles despite varying anatomies while images across clusters differ clearly; style vector dimension K=256, temperature τ=0.01, and N=8 negatives per training iteration.
Injects the style embedding via dynamic instance normalization and enables smooth style transitions between two references using spherical linear interpolation. Linear interpolation between unit vectors yields intermediate vectors of varying magnitude and inconsistent style intensity, whereas SLERP preserves unit norm and gives perceptually uniform transitions, with interpolation parameter t∈[0,1] providing explicit control over style blending. Figure 3 shows the synthesized appearance transitioning smoothly between two reference styles as t varies while anatomical content remains determined by the source input; the style consistency loss Lcon constrains the generated image to match the reference style via negative similarity.
On the CrossMoDA benchmark, style diversity yields more reliable downstream segmentation. Compared with sub-domain methods producing only 3 distinct styles, IntraStyler generates 7 style references; compared with the prior challenge-winning Dynamic method, it is significantly better on nearly all Dice and ASSD metrics. Test set of 96 T2 images; Intra-VS Dice 0.69±0.12, Extra-VS 0.81±0.12, Cochlea 0.83±0.03; ASSD Extra-VS 0.60±0.31 and Boundary 0.65±0.33; failures Intra-VS 1, Extra-VS 0, Cochlea 0; significance assessed with the Wilcoxon signed-rank test with Bonferroni correction (p<0.05).
Perspective
The result targets cross-modality MRI domain adaptation: a labeled ceT1 source domain and an unlabeled T2 target domain with substantial intra-domain variability from multiple institutes, four scanner manufacturers (Siemens, Philips, GE, Hitachi), 1.0T/1.5T/3.0T field strengths, and multiple acquisition sequences. The method suits sites where sub-domain labels are unavailable or too coarse, since each target-domain image can serve directly as a style reference at inference time and the approach scales to any degree of intra-domain heterogeneity without additional priors. The style conditioning mechanism is architecture-agnostic, and the authors note that combining it with diffusion or flow-matching models is a promising direction for higher-fidelity synthesis; the learned style embedding space shows meaningful clustering without supervision, suggesting potential use for unsupervised scanner harmonization and image retrieval.
The intensity perturbation functions for negative sampling are designed for MR and may require modality-specific tuning when moving to CT or PET, a scope limit the authors explicitly note. The global style vector from average pooling discards spatial information, so local style variations such as tumor-specific textures cannot be controlled; the authors suggest combining it with local tumor augmentation approaches. The comparison is limited to image-synthesis strategies and does not include domain adaptation approaches without explicit synthesis, such as feature-level alignment, which the authors state is beyond the scope of the study. In addition, evidence for style disentanglement comes mainly from clustering visualizations and synthesis examples, so readers may watch for validation on more modalities and larger cohorts.
