Skip to main content
Back to timeline
arXivSource publication:

Three-objective fusion and a multi-branch architecture let a wireless foundation model beat single-objective encoders on six downstream tasks with about one-third of their parameters

Related research and updates

Synopsis

The work proposes a fusion framework combining representations from three self-supervised objectives, MAE, JEPA, and SimCLR, and a multi-branch architecture with a shared trunk and objective-specific branches; after pretraining on roughly 25,000 spectrogram and CSI samples, at equal output embedding dimension fusion outperforms the evaluated single-objective encoders on all six downstream tasks spanning communication, sensing, and positioning while using about one-third of their parameters, and the multi-branch architecture retains much of the fusion benefit within a single-encoder parameter budget, with representation analyses indicating complementary contributions across objectives and much of the added benefit retained in components orthogonal to MAE's representation subspace.

Source-provided article image: Scaling of Wireless Foundation Models via Representation Diversity and Multi-Branch Architectures
Fig. 1 ·

Fig. 1 : Representation fusion framework. Three encoders, each pretrained with a distinct self-supervised objective, are frozen and their representations are pooled (mean pooling for MAE and JEPA; [CLS] token for SimCLR) and concatenated to form the fused embedding 𝐳 fused \mathbf{z}^{\text{fused}} .

arXiv

Interpretation

Three-objective fusion at an equal 768-d output outperforms the evaluated single-objective encoders on all six downstream tasks while using roughly one-third of their parameters. Prior wireless self-supervised work almost universally commits to a single objective family and combines at most two; this work fuses all three families, reconstructive, predictive, and contrastive, and frames representation diversity explicitly as a scaling axis compared head-to-head against width scaling. Pretraining on a heterogeneous corpus of about 25,000 spectrogram and CSI samples, six tasks spanning communication, sensing, and positioning, with reported gains such as 7.5 percentage points in beam prediction and a 0.60 m reduction in positioning error, compared against 256-d, 512-d, and 768-d single-objective encoders at matched output dimension.

The multi-branch architecture retains much of the fusion benefit within a single-encoder parameter budget via a shared trunk plus objective-specific branches. Fusion requires maintaining three independent encoders, tripling parameters; the multi-branch design shares early transformer blocks and specializes deeper layers, reaching a 768-d representation at zero parameter overhead, with performance stable across the tested shared-depth configurations. The zero-overhead configuration exceeds the best 768-d single-method encoder on four of six tasks with roughly one-ninth of its parameters, including a 2.0 percentage-point gain on interference classification; versus fusion, classification accuracy stays within 2.7 percentage points and positioning error increases by only 0.05 m.

Representation analyses indicate that the three objectives learn structurally distinct representations with non-redundant contributions, and that complementary information largely lies in components orthogonal to MAE's representation subspace. Prior work largely reports performance comparisons; this work provides a mechanistic account through CKA, probe weight, leave-one-out, and subspace decomposition, explaining where the fusion gain comes from. Cross-method CKA values range from 0.52 to 0.84, no probe weight share falls below 23%, leave-one-out shows removing any encoder degrades performance on at least three of six tasks, and subspace decomposition shows the orthogonal residual retains nearly all of the gain while the parallel component adds effectively nothing.

Perspective

The results apply to wireless self-supervised settings with limited pretraining data and downstream tasks spanning communication, sensing, and positioning, particularly across spectrogram and CSI modalities. For wireless systems researchers and engineers seeking to improve multi-task generalization without increasing the parameter budget, the work provides two reproducible implementation paths, the fusion framework and the multi-branch architecture, along with diagnostics such as CKA, probe weight, leave-one-out, and subspace decomposition that can be used to judge the value of objective combinations in their own tasks.

The pretraining corpus of about 25,000 samples includes proprietary spectrogram data, so whether the results change with larger data scale remains to be examined; the multi-branch model is lower than the best single-objective encoder on LoS/NLoS and jamming classification, suggesting a possible trade-off when allocating a fixed parameter budget across objectives; JEPA removal produces negligible changes on some tasks, so the task-dependent contributions of each objective warrant further study across more tasks and modalities; adaptive loss weighting and shared-branch partitioning remain open directions.

Sources