LINEUP replaces per-user private LoRA with shared low-rank factors plus eight scalars, leading on all twelve metrics across six tasks
Related research and updatesSynopsis
The work proposes LINEUP: it learns a bank of reusable low-rank personalization factors, composes them through user-conditioned recall and query-dependent calibration, and restricts target-user adaptation to a tiny user code over a shared correction space (each target user optimizes only eight scalars, versus 4.19 million per-user parameters in the evaluated private-LoRA configuration); across six tasks spanning personalized classification, prediction, and generation, LINEUP leads on all 12 metrics, each averaged over three independent runs (e.g., reducing LaMP-3 RMSE by 11.4% relative to the strongest baseline), and it maintains advantages under limited history.
Figure 1: Empirical evidence for personalization capacity allocation. (a) Small shared subspaces capture substantial held-out adapter-update energy. (b) Direction-utility rankings are substantially more consistent within users than across users. (c) A user’s own history yields positive transferable correction, unlike random or label-matched donor histories (separate y y -axis scales).
arXivInterpretation
Independent user adapters contain substantial cross-user reusable structure, the utility of reusable directions reflects both user relevance and variation across queries, and user histories provide transferable signals for compact individual correction. Reframes personalization as a capacity-allocation question: how much adaptation capacity can be shared, how shared capacity should be composed, and how much must remain user-specific, answered through three complementary empirical analyses. The three complementary empirical analyses described in the abstract, based on observations about the structure of independent user adapters; the specific datasets, models, and statistics are not given in the abstract.
LINEUP learns a bank of reusable low-rank personalization factors, composes them through user-conditioned recall and query-dependent calibration, and confines target-user adaptation to a tiny user code over a shared correction space. Decouples expressive personalization capacity from per-user trainable state: each target user optimizes only eight scalars while all shared components remain fixed, compared with 4.19 million per-user parameters in the evaluated private-LoRA configuration. The parameter-count comparison reported in the abstract (eight scalars versus 4.19 million per-user parameters) is a design-level quantification; effectiveness evidence comes from the six-task evaluation below.
The theoretical analysis gives a finite-step, finite-history risk bound and sufficient conditions for user-code refinement to improve on history initialization. Provides a risk characterization and improvement conditions for the shared-capacity-plus-tiny-user-correction design, rather than relying on empirical results alone. The theorem-style results stated in the abstract (finite-step, finite-history risk bound and sufficient conditions); proof details and assumptions are not expanded in the abstract.
Across six tasks spanning personalized classification, prediction, and generation, LINEUP leads on all 12 metrics and maintains advantages under limited history. Relative to the strongest baseline, for example reducing LaMP-3 RMSE by 11.4%; each metric is averaged over three independent runs. An evaluation over six tasks, 12 metrics, and three independent runs per metric; the abstract does not list baseline names, dataset sizes, or significance tests.
Perspective
The result targets large language model services that must personalize for many users, especially when user histories are limited and training a full adapter per person is impractical. It suggests two follow-up paths: extending the reusable low-rank factor bank and conditional composition to more task forms and modalities, and applying the shared-capacity-plus-tiny-user-correction allocation to other components that need per-user customization. The theoretical finite-step, finite-history risk bound and the sufficient conditions for user-code refinement to improve on history initialization offer a testable starting point for when refinement pays off, as a function of history length and number of steps.
This document is abstract-level material and does not include figures, dataset and baseline lists, hyperparameter settings, or proof details, so the specific values behind the 12 metrics, the composition of baselines per task, the variance across the three runs, and the assumptions underlying the theoretical results cannot be checked here. Readers who want to know under which history lengths and task distributions eight scalars remain sufficient, how the shared factor bank scales with the number of users, and when user-code refinement stops improving on history initialization will need the full experiments and proofs in the paper.
