Skip to main content
Back to timeline
arXivSource publication:

Hierarchical Spectral Shrinkage (HSS) beats separate, fully pooled, and shared-subspace estimators on synthetic experiments and multi-study gene-expression data

Related research and updates

Synopsis

The authors introduce Hierarchical Spectral Shrinkage (HSS), an empirical-Bayes framework that shrinks each task's leading spectral directions toward a data-adaptively learned common basis with shrinkage varying across tasks and spectral components, yielding the posterior mean of singular vectors in a surrogate Bayesian regression; across synthetic experiments and three immune-cell gene-expression datasets (sample sizes 156, 628, 146), HSS improves estimation accuracy and out-of-sample performance over separate, fully pooled, and shared-subspace approaches.

Source-provided article image: Empirical-Bayes spectral partial pooling across related tasks
Figure 1 ·

Figure 1 : Simulation results for the scenario (a) data-generating mechanism. Rows correspond, from top to bottom, to n 1 = n 2 = 200 , n 3 = n 4 = 500 n_{1}=n_{2}=200,\ n_{3}=n_{4}=500 , n 1 = n 2 = n 3 = n 4 = 200 n_{1}=n_{2}=n_{3}=n_{4}=200 , and n 1 = n 2 = n 3 = n 4 = 500 n_{1}=n_{2}=n_{3}=n_{4}=500 . Columns correspond to p = 500 p=500 (left) and p = 2000 p=2000 (right).

arXiv

Interpretation

HSS represents each task's leading spectral directions as a linear combination of empirical singular vectors and an aligned common basis, with a diagonal matrix controlling shrinkage separately per task and spectral component. Unlike existing grouped methods that assume common eigenvectors or a common subspace, HSS borrows strength by shrinking toward a common mean while allowing departures from the shared representation, rather than decomposing variation into shared and task-specific loading components. The estimator is derived as the posterior mean under a surrogate Bayesian regression (Eq. 9) with a hierarchical prior (Eq. 10); the shrinkage parameter is given in Eq. (17), with derivations of the marginal likelihood and coordinate updates in the supplement.

The authors provide a data-adaptive empirical-Bayes procedure that maximizes the log marginal likelihood of the surrogate regression, alternating updates of the common basis and alignment matrices, initialized by a spectral procedure and a Procrustes step. The initialization of the common basis reuses AJIVE's construction but with a different goal: finding a regularization parameter to stabilize each group's main axis of variation, without requiring equal latent dimensions across groups. The method section gives the marginal likelihood (Eq. 12), hyperparameter optimization (Eq. 13), alternating updates (Eqs. 14, 15), initialization, and shrinkage estimation (Eqs. 16, 17), with full derivations in the supplement.

Across three synthetic data-generating mechanisms, the two HSS variants are complementary: HSS-2 performs better under weak cross-group similarity or few shared directions, while HSS-1 is preferable as group eigenspaces become more similar; HSS always outperforms the no-pooling PC estimator and also the fully pooled PC estimator except when group covariances are nearly identical. Relative to the compared baselines (PC-no pool, PC-pool, PC-partial pool, AJIVE, SHARED-SUBSPACE), HSS attains lower estimation error in most settings, and its advantage persists when cross-group similarity departs from the assumed hierarchical prior. Experiments span three generating mechanisms, an unbalanced and two balanced sample-size settings, with each experiment replicated multiple times, measuring relative Frobenius error of the low-rank covariance component; some SHARED-SUBSPACE configurations had missing runs due to numerical errors, and the authors caution that those comparisons be interpreted carefully.

On three immune-cell gene-expression datasets (GSE109125, GSE37448, GSE15907; sample sizes 156, 628, 146), HSS improves out-of-sample log-likelihood over competing methods, with stronger evidence under more stringent high-variance gene filtering. The application moves the method from synthetic settings to real multi-study data, using random train-test splits and one-sided paired t-tests on out-of-sample log-likelihood rather than reporting only training error. Evidence is milder under the loosest variance filter, with several comparisons (especially HSS-2 and Groups 1 and 3) not showing strong improvement; as the filter tightens, p-values concentrate below the significance threshold for nearly all competitors and across all groups.

Perspective

The method targets settings where the same variables are measured across related but heterogeneous studies, populations, or conditions, and group covariances share similar dominant patterns while retaining differences, especially high-dimensional factor models where sample size is limited relative to dimension. It produces partially pooled estimates of group-specific loading spaces and covariance matrices and can serve as a plug-in module replacing empirical singular vectors in other SVD-based spectral estimators. Computationally, the main cost comes from initial spectral decompositions and low-dimensional coordinate-ascent updates, reducible with truncated SVD. The authors note the construction is more general and could extend to covariance and matrix-denoising estimators that retain empirical eigenvectors while regularizing eigenvalues.

The discussion explicitly lists finite-sample and asymptotic theoretical guarantees as future work, so current conclusions rest mainly on numerical experiments and one real-data application. In synthetic experiments, SHARED-SUBSPACE had missing runs in some configurations due to numerical errors, so those comparisons warrant caution. The real-data application shows milder evidence under the loosest variance filter, with some comparisons not showing strong improvement. The paper focuses on point estimation; how to use the surrogate-regression posterior to characterize uncertainty, and its frequentist coverage properties, remain open questions.

Sources