Skip to main content
Back to timeline
arXivSource publication:

FedFit cuts federated LoRA fine-tuning communication by up to 100x while matching standard federated LoRA perplexity via dual vector-bank parameterization and quantization

Synopsis

FedFit targets the large communication overhead of federated fine-tuning of large language models and the LoRA aggregation dilemma between accurate Sum-of-Products (SoP) and communication-efficient Product-of-Sums (PoS); it reconstructs high-dimensional adapter matrices from two compact, disjoint global vector banks, alternates between decoupled single-bank updates and joint updates corrected by a Residual Spectral Aggregation mechanism, and adds blockwise quantization with client-side error feedback, achieving perplexity comparable to standard federated LoRA on Qwen2.5 models with compression ratios up to 100x higher, alongside theoretical convergence guarantees.

Source-provided article image: FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization
Fig. 1 ·

Fig. 1 : Illustration of client k k in FedFit framework

arXiv

Interpretation

Introduces a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks, substantially reducing communication overhead in federated fine-tuning. Instead of transmitting full LoRA adapter matrices as in federated baselines, the transmitted object becomes compact vector banks, so communication scales with the parameterization rather than the matrix size. The abstract describes the mechanism and reports an experimental result of compression ratios up to 100x; per-configuration compression details are not expanded in the abstract.

Devises an alternating optimization schedule to resolve the SoP versus PoS aggregation dilemma, cycling between decoupled single-bank updates that allow accurate aggregation and joint updates corrected by a Residual Spectral Aggregation mechanism. Prior federated LoRA must choose between SoP aggregation accuracy and PoS communication efficiency; this schedule aims to retain both. The abstract states the mechanism and claims theoretical convergence guarantees; the specific conditions and rates are not given in the abstract.

Integrates blockwise quantization with client-side error feedback to further compress the transmitted vectors. On top of the communication reduction already obtained from vector-bank parameterization, quantization compresses further while client-side error feedback compensates quantization error. This is a method description at the abstract level; quantization bit-width and error-feedback settings are not specified in the abstract.

Experiments on Qwen2.5 models show FedFit achieves perplexity comparable to standard federated LoRA methods while providing compression ratios up to 100x higher. Presents large communication compression and non-degraded perplexity together, rather than trading accuracy for communication. The abstract reports the model family and the direction of the result, but not per-item numbers for datasets, client counts, compression ratios, or perplexity.

Perspective

The work targets settings that require privacy-preserving fine-tuning of large language models across multiple clients, especially deployments with limited communication bandwidth that still want LoRA-style low-rank adaptation. Its design goal is to shift the transmitted object from high-dimensional adapter matrices to compact vector banks while balancing aggregation accuracy and communication efficiency, so it fits federated workflows centered on adapter fine-tuning that can accommodate an alternating update schedule. The abstract's results are established on Qwen2.5 models, indicating a reusable methodological framework for similar Transformer architectures and LoRA-style adapters; for researchers and engineering teams aiming to lower cross-device communication budgets while maintaining perplexity levels, this parameterization and schedule offer a directly borrowable design.

The abstract does not state the datasets, number of clients, model sizes, or the specific configurations behind the compression ratios, so the conditions under which 'comparable perplexity' and 'up to 100x compression' hold still need checking against the full text. The assumptions underlying the theoretical convergence guarantees (such as objective properties, client sampling, and quantization error bounds) are not expanded in the abstract, and the convergence rate and conditions require reading the paper. The quantization bit-width, the implementation of error feedback, and the period settings for single-bank versus joint updates in the alternating schedule all affect the practical communication-accuracy trade-off, and the range of these hyperparameter effects remains an open question. In addition, the abstract does not report downstream task performance beyond perplexity, so whether compression affects specific task capabilities is yet to be verified.

Sources