GenFT: A Generative Parameter-Efficient Fine-Tuning Method for Pretrained Foundation Models
Synopsis
This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
Fig. 1: Comparison of LoRA and GenFT. LoRA (left) learns task-specific updates under Eq. (1). GenFT (middle) follows the conditional motivation in Eq. (2). The GenFT generator (right) produces ∆W (ℓ) = Gθ(W (ℓ)
· Page 4Interpretation
It reframes PEFT from learning ΔW largely independently to generating ΔW conditioned on W0, instantiated by a deterministic generator Gθ(W0). Prior methods exploit W0 mainly through initialization (e.g., PiSSA initializes adapters from the SVD of W0), decomposition (e.g., DoRA splits magnitude and direction), or element-wise interaction, whereas GenFT makes W0 a direct condition for generating updates throughout optimization. The paper presents the contrast between Eq. (1) and Eq. (2) as conceptual motivation and states explicitly that Eq. (2) is a conceptual motivation rather than a full probabilistic generative model, realized in practice by a deterministic generator.
The generator consists of row and column transformations that extract information from the output-channel and input-feature structure of W0, combined with nonlinear activations and masking. The paper notes that existing methods do not explicitly model row- and column-wise transformations of W0 for update generation; GenFT forms F_row = σ1(ratio·W0U)⊙Mp and F_col = σ2(F_row^T V)⊙Mp to produce ΔW. Ablations show that removing the row transformation lowers VTAB-1K average from 74.50 to 70.56 and FGVC from 90.38 to 89.96; removing the column transformation lowers them to 70.32 and 89.76 respectively.
A shared-specific decomposition splits transformation factors into cross-layer shared parts (Us, Vs) and layer-specific parts (A(ℓ), B(ℓ)), balancing cross-layer reuse and layer-wise flexibility under a small parameter budget. The paper states that this decomposition lets GenFT use a larger latent transformation dimension a+b than LoRA under a comparable parameter budget (Appendix A derives 2Da+2LDb=2LDr, giving r<a+b). Removing the shared configuration causes the largest drop in ablations: VTAB-1K average falls to 63.18 and FGVC to 84.30; Table 6 shows GenFT reaching 71.50 on Cifar with 0.26M parameters while LoRA reaches 65.79 with 3.10M.
It reports competitive average results across NLP and CV benchmarks and a feasibility pilot for generative tasks on LLaMA-7B. On GLUE, GenFT reaches 85.87% average with 0.24M parameters versus LoRA's 83.99%; on VTAB-1K it reports the highest average in the table at 74.50%; on FGVC it reports the highest average at 90.38%. On GLUE, GenFT ranks fourth on STS-B and third on QQP and QNLI; the LLaMA-7B pilot uses only a randomly selected subset of 1,000 samples and compares perplexity, and GenFT requires a larger learning rate (3×10⁻³ versus 1×10⁻⁴).
Perspective
The results target settings where a pretrained foundation model must be adapted under limited parameter and memory budgets: GLUE with RoBERTaBase, and VTAB-1K and FGVC with ViT-B/16, where after training ΔW can be materialized and merged into W0+ΔW so the extra generator cost mainly falls on training. It offers reusable components for exploring the direction of generating updates conditioned on W0 (row/column transformations, shared-specific decomposition, and tuning of the latent transformation dimension a+b), and is intended for researchers and practitioners who design PEFT methods and want stable average performance across both vision and language tasks.
The generative-model part is currently a resource-constrained pilot: it uses only a randomly selected subset of 1,000 Alpaca samples and compares perplexity, and GenFT requires a larger learning rate than LoRA and PiSSA (3×10⁻³ versus 1×10⁻⁴), which the paper attributes to its expanded rank space and describes as distinct optimization dynamics warranting larger-scale experiments. In addition, GenFT is not first on STS-B, QQP, and QNLI in GLUE, and the paper notes these tasks are sensitive to hyperparameter configurations; the latent transformation dimension a+b should also not be read as the exact algebraic rank of the final ΔW, since nonlinear activations and masking can change the resulting matrix rank. Readers may continue to watch: performance on larger-scale generative tasks, the sources of hyperparameter sensitivity, and whether the transformation differences reflected in the row/column visualizations hold across more tasks.
