Tsinghua survey on parameter-efficient fine-tuning: LoRA trains only 4.7M parameters on GPT-3, saving over 99.97% and slightly improving on full fine-tuning
Synopsis
This survey by Tsinghua University's Knowledge Engineering Group systematically reviews parameter-efficient fine-tuning (PEFT) for foundation models (FMs), organizing methods into five categories—selective, additive, prompt, reparameterization, and hybrid—and reviewing their applications across five model structures (LLMs, vision foundation models, vision-language models, visual content generation models, and multimodal foundation models), while deriving three trend observations from Semantic Scholar citation counts (PEFT growing broadly; LLMs and VFMs dominating; MFMs relatively underexplored) and identifying reliability, interpretability, and unified benchmarks as future directions.
Interpretation
The survey proposes and unifies a five-category taxonomy of PEFT: selective (e.g., BitFit, PASTA, FISH), additive (e.g., Bottleneck Adapter, MAD-X, AdapterDrop), prompt (e.g., Prefix Tuning, Prompt Tuning, P-Tuning v2), reparameterization (e.g., LoRA, QLoRA, MPO), and hybrid (e.g., UniPELT, COMPACTER, S4), with tables reporting trainable-parameter ratios for each method. Prior surveys (Xin et al., Han et al., Zhou et al., Wang et al.) focused on single directions such as visual PEFT, LLM algorithms, or methodology; this work consolidates those scattered insights into a unified taxonomy and comparison across five FM types. Tabulated summaries list each method's venue, applicable FM type, fine-tuning position, and trainable-parameter percentage, e.g., BitFit at 0.01%–0.09%, LoRA at 0.02%–0.31%, and ControlNet at 22.8%–24.7%.
Using GPT-3 as an example, the survey quantifies PEFT's cost-effectiveness: full fine-tuning involves all 175B parameters, whereas LoRA requires training only 4.7M or 37.7M, saving over 99.97% of parameters, with results showing a 0.1% to 0.5% improvement over full fine-tuning. A concrete, comparable number turns 'parameter-efficient' from a qualitative label into a verifiable order-of-magnitude comparison, and notes that PEFT is not always slightly behind full fine-tuning. The figure comes from the survey's own comparison statement about GPT-3 and LoRA, a citation-based summary of prior work rather than a new experiment in this paper.
Based on Semantic Scholar citation counts, the survey reports three trends: PEFT is growing remarkably across language, vision, and multimodal tasks; LLMs and VFMs dominate the current landscape while VLMs and VGMs gain traction as secondary areas; MFMs remain relatively underexplored, suggesting significant opportunities. It converts the impression that 'PEFT is hot' into per-model-type, per-year citation trend charts and uses them to identify MFMs as a relatively open area. The trend claims rest on Figure 1's year-by-year citation counts broken out by LLM, VFM, VLM, VGM, and MFM, i.e., bibliometric evidence.
The survey reviews applications across five FM types: for LLMs it distinguishes causal LLMs (e.g., LLaMA-adapter needing only 1.2M trainable parameters) from prefix LLMs (e.g., the P-tuning series fine-tuning ChatGLM with 0.1–0.3% of parameters); for VFMs it covers image recognition and video understanding; for VGMs it centers on diffusion models where LoRA, ControlNet, and adapter methods are most used; for MFMs it notes LoRA and Q-Former prevailing in Monkey, mPLUG-Owl, CogVLM, and GLM-4V. It carries one PEFT taxonomy across five model structures, showing how methods vary with architecture (Transformer blocks, denoising U-Net, 3D U-Net). Supported by many named methods and parameter ratios, e.g., I2V-Adapter fine-tuning only 1% of the base diffusion model's parameters and LLaMA-Adapter V2 at 0.0006%.
Perspective
The survey targets readers who need to adapt foundation models to downstream tasks under limited compute or storage, including researchers and practitioners; its taxonomy applies to language, vision, vision-language, visual generation, and multimodal model structures. Readers can use it to choose selective, additive, prompt, reparameterization, or hybrid PEFT for a given model (e.g., LLaMA, ChatGLM, ViT, Stable Diffusion, LLaVA) and to make preliminary budget estimates from the reported trainable-parameter ratios. The survey also lists cross-disciplinary work (e.g., incorporating domain knowledge and low-dimensional priors in medical imaging), continual PEFT, architecture-specific PEFT design, scaling laws of PEFT, and brain-inspired PEFT as future directions.
The loaded text is an incomplete scope: figures (Figs. 1–6) and some tables appear as placeholders or typeset fragments, so specific values and visual details cannot be fully verified, and restatements of trend statistics and some parameter ratios rely on numbers explicitly written in the body text. The survey itself notes that PEFT methods are sensitive to hyperparameters (bottleneck dimensions, rank, layer order) and that the optimal learning rate is typically much higher than for full fine-tuning; interpretability of different PEFT methods remains a challenge; and the field lacks comprehensive benchmarks, with different studies using varied datasets and task setups, leading to inconsistent performance assessment standards. These are questions readers should keep watching when adopting specific methods.
