Skip to main content
Back to timeline
Empirical Software EngineeringSource publication:

MergeRepair merges multiple code-task adapters and lifts automated program repair by 2.38% pass@1 and 4.01% pass@10 on StarCoder2 without extra training

Synopsis

The study introduces MergeRepair, which trains one LoRA adapter per task on the Python split of CommitPackFT (59,113 samples, with Bug Fixes/APR at 19.02%) for Development, Bug Fixes, Misc, Test & QA, and Improvement, merges them with weight-averaging, ties, and dare-ties under equal weights, and proposes continual merging that adds adapters sequentially; evaluated on StarCoder2 and Granite 3B code models with HumanEvalFix pass@1/pass@10, it finds that merging task-specific adapters can improve APR without additional training, reaching up to 2.38% pass@1 and 4.

Source-provided article image: MergeRepair: An Exploratory Study on Merging Task-Specific Adapters in Code LLMs for Automated Program Repair
Fig. 1

Fig. 1: An overview of MergeRepair.

· Page 7

Interpretation

Merging task-specific adapters can improve automated program repair performance without additional training. Prior model and adapter merging was mainly validated in NLP and computer vision; this work is the first to apply adapter merging to automated program repair and releases code and a replication package. Evaluated on the HumanEvalFix benchmark with pass@1 and pass@10 across two 3B models (StarCoder2 and Granite), three merging methods, and many adapter combinations, with paired t-test p-values and cliff's delta effect sizes reported; for example, merging APR with Improvement and Misc adapters on StarCoder2 yields 2.38% pass@1 and 4.01% pass@10 gains.

Merged performance depends more on which task adapters are merged than on how many are merged. The study systematically compares combinations of two to five adapters and finds that increasing or decreasing the number of adapters does not necessarily raise or lower performance, contrary to a 'more tasks is better' intuition. Across 45 RQ1 experiments, pass@1 improved over the APR adapter in 19 cases for StarCoder2 and 36 for Granite; the authors note that T3 (Misc) alone scores highest on the APR benchmark (31.40% for StarCoder2, 18.63% for Granite) and appears in all best-performing merged combinations.

In continual merging, the order in which adapters are merged significantly affects final performance, and placing the most effective adapter last improves results. The authors propose continual merging, a new paradigm that assigns geometrically decreasing weights to earlier adapters so the last-added adapter has the highest weight, simulating real projects where tasks arrive over time. RQ3 runs all 24 permutations of four tasks across three merging methods and two models; results show that adding a low-performing adapter (such as T5) last degrades performance, while adding a high-performing adapter last (T1 for StarCoder2, T4 for Granite) performs better, with most results statistically significant.

Even when the APR adapter is excluded from the merged model, the merged adapter can still match or exceed APR-specific performance. This demonstrates the generalizability of merging: adapters from non-APR tasks can transfer to APR as long as their constituent adapters score strongly on APR. RQ2 builds 11 merged adapters without the APR adapter; pass@1 exceeds the APR-specific adapter in 8/33 cases for StarCoder2 and 29/33 for Granite; for example, T2-T3 reaches 31.95% (+3.29%) on StarCoder2 and T3-T4 reaches 19.30% (+3.14%) on Granite.

Perspective

The work targets researchers and practitioners who want to reuse existing code-task adapters under limited compute and data. The experimental setting is Python, five CommitPackFT tasks (Development, Bug Fixes/APR, Misc, Test & QA, Improvement), and two 3B models (StarCoder2 and Granite), evaluated on HumanEvalFix. It shows that when tasks arrive sequentially in real projects, adapters can be merged one at a time without storing all of them simultaneously, saving memory; it also shows that in private or confidential codebase settings, merging can transfer knowledge without exposing sensitive data.

The gains are modest (up to 2.38% pass@1 and 4.01% pass@10) and depend heavily on the base model and specific task adapters; the optimal continual merging order cannot always be found by the greedy strategy based on individual performance when merging four tasks; RobustPass@1 evaluation covers only a few adapters due to compute limits, and pass@k and RobustPass@k use different input formats, so direct comparison requires caution; task labels were generated by 1-shot GPT-4 prompting and may contain noise; conclusions are limited to Python and the tested tasks, and generalization to other languages or software engineering tasks still needs verification.

Sources