Skip to main content
Back to timeline
arXivSource publication:

PerFeCT-VAR separates shared and personalized VAR dynamics via a cross-client frequency cap and attains the lowest forecast error on multi-store sales data

Related research and updates

Synopsis

The work introduces the principle of personalization diversity and PerFeCT-VAR, which decomposes each client's transition matrix into shared low-rank dynamics, shared sparse links, and personalized sparse departures, uses a cross-client frequency cap to derive a sharp shared-personalized identification threshold, and designs Frequency-Capped Thresholding (FCT) for federated estimation; the theory establishes joint linear convergence and statistical rates showing that under sufficient personalization diversity the shared dynamics retain total-sample-size gains while personalized components achieve client-level accuracy; simulations and a multi-store retail revenue application demonstrate predictive and interpretive benefits.

Source-provided article image: Personalized Federated Vector Autoregression with Personalization Diversity
Figure 1 ·

Figure 1: Three-component transition structure of PerFeCT-VAR .

arXiv

Interpretation

Introduces the principle of personalization diversity: genuinely personalized relationships should recur in only a limited fraction of clients, formalized through a cross-client frequency cap. Prior sparse personalized representations constrained only within-client sparsity and could not rule out the same relationship recurring in most clients and being absorbed into the shared component; this work makes frequency rather than magnitude the structural constraint and yields an identification condition for the shared-personalized decomposition. Proposition 1 proves via a coordinatewise majority principle that when the frequency cap is below half the clients the shared and personalized components are uniquely determined, and states the threshold is sharp for uniform identification; Appendix A.1 provides the proof and counterexample constructions.

Develops the PerFeCT-VAR estimation procedure: FCT jointly enforces clientwise sparsity and the cross-client frequency cap, combined with ScaledGD updates for the shared low-rank component, capacity-aware thresholding for the shared sparse component, and a federated initializer based on local low-rank-plus-sparse fits. Independent per-client hard thresholding may retain the same coordinate in too many clients and violate the frequency cap; FCT formulates support selection as a maximum-weight bipartite b-matching solved exactly via its linear programming relaxation, preserving the identifying structure throughout iterations. Exactness of FCT is supported by the bipartite matching and totally unimodular LP relaxation argument; Theorem 2 proves the federated initializer enters the local contraction region with high probability, and Theorem 1 gives iterative convergence.

Establishes a three-component joint contraction theory that explicitly characterizes propagation of personalized errors into shared updates, introducing directional covariance discrepancy measures and a distributional compatibility condition. Prior personalized federated analyses mostly assumed identical distributions or handled only parameter heterogeneity; this work separately quantifies how client-specific stationary covariance differences act along relevant error directions, showing covariances may differ substantially as long as their differences do not align too strongly with error directions. Theorem 1 gives a componentwise contraction inequality in which personalized-to-shared coupling terms are governed jointly by personalization frequency, sample-size imbalance, and directional covariance discrepancy; Corollary 2 gives explicit balanced-client rates where the shared rate carries the total-sample-size gain and the personalization cost is controlled by frequency.

Simulations and the Dominick's multi-store retail revenue application show predictive and interpretive benefits. Across 200 replications PerFeCT-VAR attains smaller relative error than Pooled LRS, Local LRS, Ditto, and the federated initializer, with shared-component errors decreasing in the number of clients while personalized error stays roughly flat; the communication-efficient implementation reduces uplink from 146.48 MiB to about 10.83-20.26 MiB with final errors close to the exact algorithm. On the real data PerFeCT-VAR has mean and median test MSE of 1.779 and 1.635, lower than Pooled OLS, Pooled LASSO, Pooled LRS, Ditto, the Local methods, and iAR(1); the estimated structure shows shared low-rank patterns, sparse links, and store-level adjustments.

Perspective

The framework applies when clients observe the same set of variables, each follows a stationary VAR(1), and raw series cannot be centralized, with personalized relationships satisfying the cross-client frequency cap; the theory gives explicit rates under balanced clients and Gaussian innovations, and allows stationary covariances to differ as long as a distributional compatibility condition holds. For practitioners who want shared dynamics plus store- or unit-level adjustments without sharing raw data, the work offers an implementable algorithm, communication-efficient variants, and an interpretable three-component decomposition.

The theoretical guarantees rely on stationarity, Gaussian innovations, low-rank-sparse separation, and distributional compatibility, and how well real data satisfy these conditions still needs case-by-case checking; how to choose the frequency cap, rank, and sparsity levels in real applications is handled by validation-set MSE minimization, but sensitivity is not expanded in the main text; gradient compression and less frequent FCT refresh are approximate operations whose error behavior under larger scale or stronger heterogeneity merits further observation; the real-data experiment is a single supermarket-chain dataset, so generalization to other industries and longer forecast horizons remains an open question.

Sources