FedLore shares a low-rank gradient basis within each round and refreshes it across rounds, beating the evaluated low-rank adapter baselines and matching or exceeding full-parameter training on vision and language tasks while cutting communication and optimizer-state memory
Synopsis
The work proposes FedLore, which shares a low-rank optimization basis within each round and refreshes it across rounds, enabling exact aggregation in low-rank coordinates, eliminating the identified projection bias, and letting the accumulated model update exceed the per-round rank budget; the authors characterize the aggregation bias and establish an O(T^{-1/2}) stationarity bound for the projected-SGD variant under a global-gradient coverage condition with standard smoothness and variance assumptions and bounded gradient heterogeneity; on vision and language tasks, including federated pre-training, FedLore outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training while reducing communication and optimizer-state memory.
Figure 2: Fixed-rank LoRA vs ours low-rank updates. FedLore shares a low-rank update basis within each round and refreshes it across rounds, allowing the accumulated model change to exceed the per-round rank.
arXivInterpretation
The paper identifies and names subspace fragmentation: when clients independently choose low-rank subspaces, local projections interact with data heterogeneity to bias aggregated directions, and aggregation can increase update rank and communication cost, so accurate local gradient compression need not preserve global descent. Relative to LoRA-based low-rank adapter methods, whose fixed rank budget can limit adaptation, and to prior gradient low-rank optimization with independently chosen client subspaces, the work locates the issue in client subspace inconsistency rather than in per-client compression accuracy alone. The problem is presented as a conceptual characterization together with an aggregation-bias analysis, which the abstract states the authors characterize; concrete experimental scale and numbers are not given in the abstract.
FedLore shares a low-rank optimization basis within each round and refreshes it across rounds, which makes aggregation in low-rank coordinates exact and eliminates the identified projection bias, while subspace refresh allows the accumulated model update to exceed the per-round rank budget. The shared basis couples exact aggregation with low-rank communication, and cross-round refresh relaxes the single-round fixed rank budget on cumulative adaptation capacity. The mechanism is stated explicitly in the abstract; its effect is supported by experiments on vision and language tasks including federated pre-training, though the abstract lists no specific datasets, model scales, or numbers.
The authors establish an O(T^{-1/2}) stationarity bound for the projected-SGD variant under a global-gradient coverage condition, standard smoothness and variance assumptions, and bounded gradient heterogeneity. The result provides a convergence characterization for projected federated optimization under a shared basis, complementing the aggregation-bias analysis. This is a theoretical result; the abstract gives the rate and assumptions, while proof details and constants are not presented in the abstract.
On vision and language tasks, including federated pre-training, FedLore outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training while reducing communication and optimizer-state memory. Relative to low-rank adapter baselines it reduces communication and optimizer-state memory while maintaining or improving accuracy; relative to full-parameter training it does not sacrifice accuracy. The abstract reports comparisons across vision and language tasks and federated pre-training, but gives no specific metric values, baseline list, or ablation scale.
Perspective
The result targets federated foundation-model training constrained by client memory and communication costs, and applies to federated pipelines using low-rank gradient optimization and projected-SGD-style updates; it is relevant to researchers and practitioners who want to cut communication and optimizer-state memory on vision and language tasks, including federated pre-training, without losing accuracy relative to full-parameter training. The theoretical guarantee's scope is delimited by the global-gradient coverage condition, standard smoothness and variance assumptions, and bounded gradient heterogeneity.
Readers should still watch how the shared basis is chosen and refreshed under strong heterogeneity or frequent participant churn; how the global-gradient coverage condition is met in practical federated settings; how the cumulative update-rank growth from cross-round refresh affects communication and memory in practice; and the specific datasets, model scales, baseline configurations, and numeric results absent from the abstract, which require the main text tables and experimental details. Because this assessment is based on the abstract only, figures and the full experimental setup could not be checked.
