PEARL generates personalized images from user histories via a reason-reflect loop, improving personalization metrics by 15% on its PMH-IG benchmark
Synopsis
The work introduces PMH-IG, the first unified benchmark for personalized image generation from realistic user histories (Amazon reviews and Instagram posts), comprising Personalized Scene Generation and Personalized Creative Generation with a multi-axis evaluation protocol, and proposes PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward, achieving an average improvement of 15% across personalization metrics on both tasks.
Interpretation
The paper introduces PMH-IG, a benchmark for personalized image generation from realistic user histories, with two complementary tasks: Personalized Scene Generation (based on Amazon review histories, placing a given product in a scene reflecting the user's lifestyle and preferences) and Personalized Creative Generation (based on Instagram post histories, generating a novel image on a specified topic faithful to the user's aesthetic and visual identity). Existing benchmarks (e.g., Xu et al., Dunlop et al.) typically provide a few curated reference images per user and evaluate concept reproduction under novel prompts; PMH-IG instead uses naturally occurring multimodal histories and evaluates whether a model can infer the user's identity and synthesize a new image reflecting it, explicitly separating the generation condition (what to depict) from the user history (how to depict it). The paper reports dataset statistics: 5,000 users and 50,000 reviews for scene generation, and 4,413 users, 50,656 posts, and 1,271 (user, target) pairs for creative generation; splits are user-disjoint at roughly 80:10:10, so the test set directly probes generalization to unseen users.
The paper proposes PEARL: a Stage-1 planner reasons over user history to produce a chain-of-thought and scene description rendered by a frozen generator into an initial image; a Stage-2 reflector compares that image against the history, identifies personalization mismatches such as scene incongruities, missing lifestyle cues, or aesthetic deviations, and revises the plan for re-rendering. Prior history-based personalization methods (e.g., Pigeon, PMG) fuse the history directly into the generator as a unified conditioning signal without explicit reasoning about its implications; PEARL externalizes that inference as a textual plan and then verifies and corrects the plan against rendered evidence. Both planner and reflector are initialized from Qwen2.5-VL-7B, the renderer is SDXL-1.0 at 1024x1024, and LoRA adapters are used; Stage 1 distills silver trajectories from a Gemini2.5-Flash teacher with privileged access to an identity anchor, and Stage 2 uses alternating-policy DPO where each policy is scored by the full downstream pipeline including the other policy and the renderer, with all experiments run on a single 4090 GPU.
On both PMH-IG tasks, PEARL achieves an average improvement of 15% across personalization metrics and outperforms baselines including PMG, Pigeon, LaVIT, and LLaVA on most metrics. The paper highlights the value of multi-axis evaluation: in creative generation, methods that win LPIPS or MS-SSIM score poorly on retrieval, showing that a single fidelity metric can be won without delivering personalization. For scene generation, PEARL reaches H@5 of 0.2297 and MRR of 0.1639, above all baselines; for creative generation, PEARL is best on 7 of 10 columns including all retrieval and MLLM judge metrics, with CIS 0.645, DIS 0.289, inter-category R@1 0.876, and intra-category R@1 0.701; all baselines share the same SDXL renderer for fair comparison.
Ablations show the reflection module contributes measurable gains, concentrated on fine-grained personalization details. The paper compares full PEARL against PEARL-Reflection, which replaces the alternating-policy DPO stage with supervised fine-tuning alone, indicating that closing the loop between reasoning and rendering through preference optimization yields gains over supervised distillation alone. The paper reports PEARL outperforms PEARL-Reflection on four of five aggregated dimensions across both tasks and the MLLM judge, with the largest improvements on style and content; a marginal drop on the overall score (under a stated threshold) leads the authors to conclude the reflector mainly aligns fine-grained details rather than making drastic adjustments, consistent with the findings of IRG.
Perspective
The results target research and engineering settings where identity is inferred from accumulated user history to generate images, applying to e-commerce product presentation and social media content creation, and transferable to the settings listed in the appendix: personalized advertising, virtual try-on and lifestyle staging, creator tooling, personalized storyboards, generative recommendation, and evaluation of personalized multimodal agents. The benchmark and evaluation protocol are independently useful: the identity element extraction pipeline serves as a diagnostic for user identity modeling, the retrieval-based metrics support generative recommendation research, and the multi-axis suite offers a standardized testbed for personalized multimodal agents. For practitioners, the paper suggests using retrieval and recommendation metrics as a reproducible offline proxy to benchmark generative models for personalized product imagery before deployment.
Personalized Scene Generation has no naturally occurring ground-truth scene image per (user, product) pair, so user alignment is measured indirectly through retrieval-based and judge-based metrics rather than direct reference-based comparison. PEARL's training performs only a single round of alternating-policy DPO updates between planner and reflector, and the paper leaves a systematic study of multi-round optimization to future work. The human evaluation is a preliminary single-annotator judgment of 25 pairs per task, measuring an outside observer's assessment of history fit; the paper explicitly states it does not establish inter-annotator reliability, population-level preference, or statistical significance, and the creative comparison used each pipeline's respective history subset rather than identical conditioning. In addition, one scene account was subsequently flagged for target-review leakage in a separate rerendering run, its original SDXL generation inputs are unavailable for verification, and contamination of the original outputs is unresolved, with a sensitivity analysis excluding that account reported as a check. A reader who only sees the abstract would miss these scope and open questions.
