Skip to main content
Back to timeline
arXivSource publication:

AutoAdapt trains independent LoRA adapters via automatic latent domain discovery, matching an all-domain single adapter across 14 benchmarks while enabling retraining-free extension

Synopsis

AutoAdapt is a modular instruction-tuning framework that first discovers latent domains with three methods (K-means, BERTopic, MNLI), then trains one LoRA adapter per domain independently and in parallel, and routes at inference with a parameter-free centroid mechanism; across 14 domain-specific benchmarks and GPT-4o pairwise judgements it reaches parity with a single LoRA adapter trained on all domains, while adding or updating a domain requires training only one adapter instead of full-model retraining.

Source-provided article image: AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility
Figure 1 ·

Figure 1: Domain overlap between discovery methods on LLM-as-a-judge evaluation with GPT-4o.

arXiv

Interpretation

AutoAdapt partitions heterogeneous instruction data into automatically discovered latent domains, trains one independent LoRA adapter per domain, and routes at inference by cosine similarity to domain centroids, adding no trainable routing parameters. Unlike Mixture-of-Experts architectures that jointly train experts and routing, and unlike prior multi-adapter routing that needs a trained selector adapter or LLM-generated routing signals, this work extends parameter-free routing to the instruction-tuning setting and adopts a taxonomy-free domain discovery step. Trained on Mistral-7B-Instruct-v0.3 with 509,174 instructions from eight public datasets, compared against a single LoRA adapter trained on all domains, evaluated on 14 domain-specific benchmarks and blind GPT-4o pairwise preference judgements.

The three independent discovery methods reach overall parity with the single-LoRA baseline: adapter win rates against single LoRA in blind GPT-4o judging were 50.5% for K-means (858/1,700), 50.2% for BERTopic (1,205/2,400), and 46.9% for MNLI (1,004/2,143). Prior fixed-domain specialisation work showed parameter-efficient methods can match full-model fine-tuning but did not provide empirical evidence of extensibility; this work supplies both a cross-method comparison and a post-deployment extension demonstration. For each domain, 100 test instances were sampled for blind, reference-grounded pairwise comparison; on the benchmark side, accuracy differences across 14 benchmarks were combined as an inverse-variance weighted mean with 95% confidence intervals.

Discovery methods partially converge: K-means and BERTopic show an Adjusted Rand Index of 0.42 and Normalised Mutual Information of 0.55, whereas MNLI versus the others shows ARI of only 0.08/0.07 and NMI of 0.26/0.28, indicating entailment-based groupings are almost completely different. This turns the question of whether automatic discovery is stable from a single method's internal metrics into a quantified cross-method agreement comparison, and shows MNLI is the only AutoAdapt method to win on RACE, MBPP, CRUXEval and HellaSwag. ARI and NMI are computed over the domain partitions each method produces on the same data, interpreted alongside each method's win/loss pattern across the 14 benchmarks.

Extensibility is a structural property: after adding 10,000 Magicoder examples, 99.9% (8,494/8,500) of routing assignments were unchanged and the updated c15 adapter achieved the lowest NLL of 0.1926 and perplexity of 1.212; after adding a legal domain, routing stability was 99.7% (8,472/8,500) and the legal adapter was preferred in 126/200 and 129/200 comparisons on 200 held-out LawInstruct QA examples. Compared with monolithic fine-tuning that requires retraining on the combined dataset, this work directly demonstrates extension through training a single adapter in two simulated post-deployment scenarios, leaving other adapters unmodified. Both simulations report routing stability, held-out language-modelling metrics (NLL, PPL) and blind GPT-4o pairwise preferences; parallel training ran three adapters simultaneously on a single A100 80GB GPU in 10h 8min wall-clock, 1.7x faster than single LoRA and 2.0x faster than full model fine-tuning.

Perspective

The result is aimed at heterogeneous, continuously evolving instruction-tuning deployments, especially teams with limited compute that need to add or update domains frequently. It enables new domains to be incorporated by training one additional adapter and existing domains to be updated by retraining only the relevant adapter, leaving other components unchanged and avoiding monolithic retraining; parallel training ran three adapters simultaneously on a single A100 80GB GPU in 10h 8min wall-clock, 1.7x faster than single LoRA and 2.0x faster than full model fine-tuning. At inference, centroid routing replaces neural routing, with the base model resident in GPU memory and only the active lightweight adapter swapped per request. The authors state that future work should extend beyond a low-resource deployment scenario and investigate the impact on larger models.

Several open questions remain for a careful reader: discovery methods do not agree with each other, with MNLI versus the others showing ARI of only 0.07-0.08, so what 'automatic discovery' yields depends on the chosen method; coding adapters underperform single LoRA in both benchmarks and blind judging, which the authors suggest reflects that diverse coding domains benefit from shared knowledge representation; MultiNLI shows a consistent negative effect across methods, which the authors attribute to topic-based fragmentation of topically unconstrained content and propose enriching topic embeddings with functional embeddings as future work. In addition, the extensibility conclusions come from two simulated scenarios, the legal-domain evaluation uses only 200 held-out examples, and the authors explicitly scope the findings to a low-resource deployment scenario and the current base-model scale.

Sources