MIRROR turns LLM personalization from imitation to preference internalization via reference-revealed on-policy self-distillation, leading three benchmarks with less forgetting
Synopsis
MIRROR treats a user's reference text as hindsight information and aligns the model's next-token distributions along its own on-policy trajectories with those of its reference-conditioned self, internalizing preferences instead of copying wording; MIRROR-F adds selective anchoring on informative, context-grounded reference tokens. Across three personalized generation benchmarks, two model scales, and reference-based plus LLM-judge evaluations, both achieve leading personalization and text quality with less forgetting than SFT baselines on three unseen tasks.
Figure 1: Overview of personalization paradigms. (a) A case study of the quality defects of existing methods. (b) Retrieval-based personalization. (c) SFT-based personalization. (d) MIRROR, which aligns a self-student with a self-teacher on the student’s own on-policy trajectories.
arXivInterpretation
MIRROR reformulates personalized generation as reference-revealed self-distillation: one model acts as both self-teacher and self-student, where the teacher conditions on the user reference to produce preference-aware next-token distributions and the student aligns to them on prefixes it samples itself. Unlike SFT-style fine-tuning whose target is the reference token evaluated on reference prefixes, the target here is a distribution evaluated on the model's own on-policy states, so training and inference share the same state distribution. The paper specifies a stopped-gradient surrogate objective in which sampled prefixes and the teacher distribution are detached, so gradients flow only through the student distribution on fixed on-policy states; the authors argue this yields preference-reasoning behavior that transfers to queries without references.
MIRROR-F adds a selective anchoring loss applied only to reference words that are informative by deterministic surface-form rules and grounded as complete words in the student context, with strength controlled by a single focality coefficient. Rather than uniformly supervising the whole reference, the auxiliary signal concentrates on content-bearing positions such as terminology, named entities, and quantities, while the distributional term continues to carry personalized style. Selection is a word-level heuristic plus a lexical grounding check, expanded to all subword units; the objective reduces to MIRROR when the selected set is empty, and the authors report personalization metrics remain stable as the focality coefficient varies from 0.1 to 0.9.
On three personalized generation tasks from LaMP and LongLaMP with Qwen3-1.7B and Qwen3-4B, MIRROR and MIRROR-F are leading or competitive on ROUGE-1, METEOR, BERTScore, and LLM-judged content-and-style scores. Against retrieval baselines RAG, LatestK, LLM-TRSR and PEFT baselines SFT, OPPU, PerCE, NextQuill, the authors report MIRROR-F attains the best results in most settings, including ROUGE-1 and METEOR on all three tasks with the Qwen3-1.7B backbone. Each reference-based and LLM-judge evaluation uses three independent training runs with different random seeds averaged arithmetically, with Qwen3-30B-A3B as judge at fixed zero decoding temperature and 100 common samples per task.
On three unseen personalized generation tasks, MIRROR and MIRROR-F degrade less than SFT-based methods and score higher on G-Eval Coherence, Consistency, Fluency, and Relevance than SFT, PerCE, and NextQuill. The authors attribute the larger degradation of SFT-style methods on unseen tasks to their reliance on reference-specific token-level supervision, whereas on-policy distributional alignment transfers preference-relevant behavior rather than reference wording. The OOD conclusion rests on degradation magnitudes across held-out Amazon Reviews and Music Review tasks, and the text-quality conclusion on reference-free G-Eval with SummEval dimensions.
Perspective
The work targets personalized text generation where user history and user-written references exist: the teacher sees the reference during training, while the deployed student uses only user history and the current question, so results apply to settings with available references and task forms close to LaMP and LongLaMP. Training uses LoRA on Qwen3-1.7B and Qwen3-4B for 200 and 250 steps respectively, with the focality coefficient fixed at 0.5 in the main comparison, which bounds the current conclusions. For teams aiming to reduce catastrophic forgetting while keeping content professional, the framework offers a reusable objective interface: the distributional term carries style and the selective anchoring term carries content, and the anchoring rules can be swapped per domain without changing the distillation mechanism.
The mixing coefficient is fixed at 0.5, and the authors state they do not claim it is optimal, with systematic study of alternatives beyond scope; the focality coefficient keeps personalization metrics stable from 0.1 to 0.9, but anchoring strength and professional quality relate non-monotonically, with moderate values best. The rule-based information-completeness metric measures exact coverage of predefined grounded information words and does not capture paraphrase, factual entailment, fluency, or style, so it should be read jointly with G-Eval results. On Qwen3-4B news, MIRROR exceeds MIRROR-F in information completeness, indicating anchoring benefits are task- and backbone-dependent. The thinking-mode ablation shows effects varying by task and scale, leaving open how personalization and long-horizon reasoning interact.
