CARD pairs cluster-level LoRA with decoding-time preference vectors, taking 10 of 12 metric settings across six LaMP and LongLaMP tasks
Synopsis
CARD introduces a hierarchical framework for personalized text generation: it clusters users by shared stylistic patterns and trains group-specific LoRA adapters, derives lightweight user preference vectors through implicit preference learning that contrasts user-authored text with cluster-level generations, and injects personalization at inference only via low-rank logit corrections, ranking first in 10 of 12 settings across six LaMP and LongLaMP tasks and two metrics while remaining stable for low-resource users, across model scales, and in storage efficiency.
Interpretation
CARD splits personalization into two layers, a shared cluster-level prior and an ultra-lightweight user-level vector, where the cluster LoRA provides generalization and a low-resource fallback and the user preference vector carries fine-grained stylistic deviation. Prior PEFT methods typically maintain per-user parameters, which becomes expensive as the user base grows, while RAG methods prepend user history to the prompt and yield shallower personalization sensitive to retrieval quality. CARD amortizes adaptation across clusters and pushes individual differences to decoding time. Across six tasks on LaMP and LongLaMP and two metrics, ROUGE-1 and ROUGE-L, CARD ranks first in 10 of 12 settings and is near-best in the other two (LaMP5 R-1 at 0.459 vs. 0.464; LongLaMP2 R-1 at 0.252 vs. 0.255); the ablation shows removing the user vector drops LaMP-4 R-1 from 0.218 to 0.148.
Implicit preference learning builds preference pairs from user-authored text versus cluster-LoRA generations, turning stylistic deviation under matched semantics into supervision without manual annotation. The paper argues explicit preference annotation is prohibitively expensive and heuristic constructions with random negatives entangle topical content with stylistic traits; CARD's negative comes from the same cluster LoRA on the same prompt, aligning semantics while differing in style and isolating pure stylistic deviation. The preference-pair construction is specified in the method section, training uses a Bradley-Terry pairwise loss with the cluster LoRA and backbone frozen, and the paper reports the design mitigates preference-data scarcity and remains effective for low-resource users.
Inference-time personalization only edits logits: the user preference vector modulates hidden states channel-wise, a low-rank vocabulary mapping corrects the cluster-LoRA logits, and the correction is applied only to Top-k candidate tokens. Compared with per-user fine-tuning or long-context retrieval, CARD keeps the backbone and cluster parameters frozen at inference, so switching users only swaps a compact preference vector, which the paper says can be stored on the user device for on-device personalization. The paper states the Top-k restriction reduces complexity from the full vocabulary to the candidate set, reports that a new user needs only a lightweight, training-free profile encoding for cluster assignment, and presents an efficiency table covering training time, per-query latency, and per-user storage.
Evaluation combines automatic metrics, GPT-5.2 as an LLM judge, and human evaluation, which broadly agree but are not perfectly matched. The paper reports LLM and human judgments largely agree on ranking but diverge more on highly stylistic tasks; CARD improves over the non-personalized baseline by 76.4%, 95.5%, and 113.5% in LLM scores on LaMP-4, LaMP-5, and LaMP-7, with substantial human-evaluation gains as well. The appendix reports average Pearson 0.687, Spearman 0.670, Kendall 0.529, standard Cohen's Kappa 0.384, and weighted QWK 0.618, with the paper explaining that standard Kappa is low because it heavily penalizes one-point deviations on a five-point scale.
Perspective
This work targets text generation tasks where a user's writing history is available for modeling, such as news headlines, scholarly titles, tweet paraphrasing, abstract generation, topic writing, and product reviews; the paper validates on LaMP and LongLaMP with a Qwen3-8B backbone and additionally sweeps Qwen3 scales from 0.6B to 32B. For teams that want personalization without per-user fine-tuning or long histories in context, CARD offers a path of cluster-level adaptation plus decoding-time vectors, where a new user needs only a training-free profile encoding for cluster assignment and preference-vector estimation. The paper also notes the user preference vector can be stored on the user device for on-device inference while the provider releases only a lightweight personalization module, reducing exposure of raw history.
The paper's own stated boundaries include: offline grouping uses unsupervised K-means, which may not fully capture complex user relationships or latent personalization structure; each user is represented by a single preference vector at decoding time, which may limit expressiveness for diverse or evolving preferences, and the learned dimensions are not directly interpretable; the multi-stage pipeline still requires coordination across several components; and noisy or weakly relevant histories can degrade the learned user vector. The paper also notes performance drops on LaMP-7 with long histories, likely reflecting noisy histories, and leaves history filtering, relevance weighting, and noise-robust profile selection to future work. In addition, experiments use the validation set as the test set because the official test set is not public, which affects direct comparability with public leaderboards. Readers interested in absolute levels on a specific task should return to the main tables and appendix to check each task's data scale and truncation settings.
