IMaX swaps the marginal entropy term of mutual information for a Tsallis α-entropy, lifting accuracy by up to 7.3 points on ESCA and retinal datasets under long-tailed semi-supervised domain generalization
Synopsis
The work shows that semi-supervised domain generalization methods such as FBCSA and DGWM degrade substantially under long-tailed class distributions, and proposes IMaX: maximizing the mutual information between learned features and latent labels under supervision constraints from labeled samples, while replacing the standard marginal entropy term with a Tsallis α-entropy to relax the class-balance assumption; across two modalities (ESCA histopathology and diabetic retinopathy grading), three SSL frameworks and three SSDG methods, IMaX improves accuracy in all but one setting, with gains up to 7.3 points in the low-label regime (mL=5).
Interpretation
The paper introduces the long-tailed semi-supervised domain generalization (long-tailed SSDG) setting: each source domain has few labeled samples and many unlabeled ones, and the labeled class distribution follows an exponentially decaying long tail. Prior SSDG work (FBCSA, DGWM) assumes uniform class distributions across source domains; the paper shows empirically that these methods degrade substantially under long-tailed distributions, making class imbalance an explicit part of the SSDG problem definition. The paper reports accuracy comparisons between uniform and long-tailed training distributions on ESCA and Retina (Fig. 1, left), and states that imbalance is controlled by a hyperparameter γ set to 10 in the experiments.
The paper proposes the IMaX objective: maximizing mutual information subject to yi=pi on labeled samples, and replacing the unsupervised conditional entropy term with a pseudo cross-entropy where pseudo-labels from weak augmentations guide predictions on strong augmentations, yielding a semi-supervised view of mutual information. Mutual information maximization has mostly been used in deep clustering and few-shot learning; the paper adapts it to SSDG with explicit supervision constraints and notes that directly optimizing conditional entropy can lead to degenerate solutions mapping all samples to one class. The paper provides a step-by-step derivation from Eq. 2 to Eq. 6, and an ESCA ablation shows that adding only this semi-supervised mutual information term raises accuracy from 61.0 to 66.0 (mL=5) and from 69.2 to 72.6 (mL=10).
The paper replaces the standard marginal entropy H(Y) with a Tsallis α-entropy Hα(Y), writing marginal entropy as a KL divergence from the uniform distribution and generalizing it to an α-divergence form, thereby relaxing the strong bias toward class balance. Standard marginal entropy pushes the predicted marginal distribution toward uniform, implicitly assuming class balance; the α-entropy tolerates marginals further from uniform and reduces to the original form when α=1. The ESCA ablation shows that swapping in Hα on top of the semi-supervised mutual information term further raises accuracy to 68.3 (mL=5) and 73.1 (mL=10); α is selected on the validation set (1.5 for ESCA, 2 for Retina), and Fig. 2 shows validation and test accuracy follow highly consistent trends across α.
IMaX is model-agnostic and plug-and-play, stacking onto three frameworks (Baseline, FBCSA, DGWM) and three SSL methods (FixMatch, FreeMatch, StyleMatch), and generally improving accuracy across the ESCA and Retina modalities. The paper emphasizes that the method does not rely on modality-specific assumptions, in contrast to medical-imaging-specific approaches that assume shifts in hue, saturation and contrast or preserved structural similarity across domains, giving it broader applicability. Table 1 covers 3 frameworks × 3 SSL methods × 2 datasets × 2 label budgets, i.e. 36 settings, of which 35 improve and one declines (FBCSA+FixMatch on Retina, mL=10, 32.0→29.5); each experiment consists of 20 runs (four leave-one-domain-out setups × five seeds), with the largest gain of +7.3 in the low-label mL=5 regime.
Perspective
The result targets semi-supervised domain generalization where each source domain has few labeled samples and many unlabeled ones, and the labeled class distribution is long-tailed, with the model deployed directly to an unseen target domain different from all source domains. The paper validates on two modalities: histopathology (ESCA, 11 classes, four hospitals as domains) and ophthalmology diabetic retinopathy grading (Messidor-2, IDRiD, Paraguay, APTOS as four domains), with α selected on the validation set (1.5 for ESCA, 2 for Retina) and imbalance controlled by γ=10. IMaX can therefore be attached as an additional objective to pseudo-label-based SSL/SSDG training pipelines, for researchers and practitioners seeking cross-domain generalization under scarce labels and imbalanced classes.
The paper reports classification accuracy on two image modalities and does not cover segmentation, detection or non-image tasks, so whether those directions benefit similarly remains open. In Table 1, FBCSA+FixMatch on Retina with mL=10 shows a decline from 32.0 to 29.5, indicating gains are not uniform across all framework and data combinations, and the conditions for that are worth watching. The value of α depends on the validation set; the paper shows validation and test trends align, but this transferability still needs checking when validation and target domains differ more. In addition, the paper uses γ=10 as the main imbalance setting, and although Fig. 1 (right) shows trends as the imbalance factor varies, behavior under more extreme imbalance remains an open question.
