Skip to main content
Back to timeline
arXivSource publication:

PsPLUG uses a "personalization residual" plug-in to curb style-instruction suppression of user preferences, outperforming baselines on LaMP

Synopsis

The work identifies that explicit style instructions can erode the user-specific traits that personalized LLMs aim to preserve (a failure mode it calls personalization collapse), proposes modeling personalization as a distributional residual between the user's true linguistic distribution and the base model's neutral distribution under the same input, and builds PsPLUG: a lightweight plug-in that prepends a 3-token prefix (system instruction vector, user vector, input vector) to a frozen Qwen3-8B backbone, learns the user-specific residual with a style-conditioned Bradley–Terry preference objective, and tunes personalization strength at inference via a scaling coefficient; on the LaMP benchmark, PsPLUG generally outperforms non-personalized, RAG, PAG, PPlug, and OPPU baselines without styl

AI-generated editorial illustration: Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs

Interpretation

The paper identifies and names a failure mode: explicit style control can conflict with implicit user preferences, producing "personalization collapse" in which strong system instructions dominate the generation space and override diverse user-specific traits. Prior personalization work focused mainly on injecting user signals (retrieval-based, per-user fine-tuning, soft-prompt plug-ins) and rarely characterized how those signals relate to the base model's neutral behavior under the same input, nor systematically examined interference from style instructions. The finding comes from the authors' experiments on LaMP text-generation tasks with four predefined style instructions (Warm, Critical, Concise, Elaborative); the paper also notes that LaMP relies mainly on ROUGE-1/ROUGE-L surface overlap, making the style–personalization interaction "inherently entangled and difficult to disentangle," so it adds LLM-judge and human style scores.

The paper proposes viewing personalization as a "distributional residual": the deviation of the user's true linguistic distribution from the base LLM's neutral distribution under identical input, isolated via a style-conditioned preference objective. Unlike standard preference optimization that merely anchors to generic base distributions (e.g., LongPO as cited), this residual reward mathematically disentangles the two competing signals to preserve fine-grained persona fidelity under strict stylistic constraints. The method is implemented with a Bradley–Terry pairwise loss, where preference pairs contrast user-authored text against a style-only baseline that follows the instruction but contains no user-specific information; without style instructions the framework reduces to a special case referencing a neutral baseline.

The paper delivers PsPLUG as a lightweight plug-in: a 3-token prefix (system instruction vector, user vector, input vector) injected into a frozen backbone, where the user vector comes from a PAG text profile of user history encoded by a frozen sentence encoder (BGE-base-en-v1.5) and mapped by a trainable MLP, and the input vector comes from pooled frozen-backbone features mapped by an MLP. Because the backbone stays frozen and user signals are decoupled into a modular prefix, users can be switched rapidly without reloading model parameters, and a scaling coefficient on the user vector provides continuous inference-time strength control. Implementation details are explicit: Qwen/Qwen3-8B backbone, prefix injected at the input embedding layer, prefix length 3 tokens, 5 training epochs with early stopping, bfloat16 precision, 8 NVIDIA H100 GPUs; the efficiency comparison reports that RAG storage scales with history size and adds latency, PEFT requires per-user optimization and adapter-loading latency, whereas PsPLUG encodes the user profile once into a single static embedding, keeps prompt length constant, and makes compute independent of user history size.

Experiments show PsPLUG generally outperforms baselines without style instructions, preserves personalization under style constraints while achieving the highest style scores, and offers a tunable trade-off between personalization and style adherence via the strength coefficient. The paper reports best or second-best results in over 80% of style–task–metric settings, with gains up to about 2–4 ROUGE points over prior baselines, and the highest style scores across all four styles. Main results cover six LaMP tasks (LaMP-1/2M/3/4/5/7) with their metrics (ACC, F1, MAE, RMSE, ROUGE-1, ROUGE-L, METEOR); style experiments cover LaMP-4, LaMP-5, and LaMP-7 across four styles; style and persona scores come from an LLM judge with human validation, and the appendix reports system-level correlations (style: Pearson 0.858, Spearman 0.864, Kendall 0.725; persona: Pearson 0.704, Spearman 0.712, Kendall 0.589).

Perspective

The result targets personalized generation settings that must preserve user preferences under explicit style instructions, such as news headline generation, scholarly title generation, and tweet paraphrasing; its users are engineering and research teams wanting low-overhead personalization with on-demand strength adjustment. The method assumes access to user history and offline profile embeddings, keeps the backbone frozen, and thus suits multi-user switching with constant prefix length; the paper positions the framework as "a preliminary step" that poses the question of disentangling personalization from style.

Style conditions currently cover only four predefined styles (Warm, Critical, Concise, Elaborative); finer-grained, open-ended, and compositional real-world style demands are not yet included. Experiments are limited to specific backbones within the LaMP benchmark distribution, leaving robustness across architectures, multilingual settings, and long-context domains an open question. On evaluation, LaMP relies mainly on surface-overlap metrics such as ROUGE-1/ROUGE-L, so the style–personalization interaction cannot be fully separated; the paper therefore adds LLM and human judgments, but the two diverge more under Critical and especially Elaborative styles, indicating that judging fine-grained persona cues remains subjective. In addition, this evidence bundle merges the full paper with an external story, and some tables and figures (e.g., the strength-sensitivity figure and case-study figure) appear as references rather than complete values, so the specific strength-scaling curves can only be understood from the main text.

Sources