UMM-Reflection lifts BAGEL's GenEval from 0.72 to 0.84 and its repair rate from 20.59% to 64.94% by training on whole reflection trajectories
Synopsis
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
Interpretation
It makes multi-round reflection a trainable whole: sibling trajectories sampled from one shared initial image let a group-relative advantage compare reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Prior unified-model work either applied RL to a single render (T2I-R1, ReasonGen-R1, UniRL) or imitated multi-round reasoning-and-editing trajectories (Thinking with Generated Images, MINT, Uni-CoT, IRG, ThinkMorph, UniT); none applied outcome-driven RL to the model's own multi-round reflection. The paper provides the derivation plus ablations: training only the flow head stays near SFT (73, repair rate 22.8%), training only the text head recovers most of the gain (78, 49.4%), and joint training reaches 84 and 64.9%, using a 3,000-prompt pool, two roots and sixteen siblings per root, and 1,000 updates.
It separates producing useful revisions from reliably choosing them: SFT already learns the protocol (95% of trajectories comply) and its rollouts already contain correct repairs, yet a single SFT trajectory repairs only 20.59% of initially incorrect images, rising to 64.94% after RL. It carries SCoRe's argument that online multi-turn RL is needed for self-correction from language into image reflection inside a unified model, with quantified evidence. On the 553 GenEval prompts RL beats SFT on 109 and loses on 38 (McNemar), while on initial images the split is 48-32 and not significant, locating the gain in the reflection rounds; training logs show pass@1 rising from 22.7% to 70.1% while pass@16 rises only from 77.6% to 97.1%.
The gains transfer and no verifier is needed at inference: the same checkpoint improves over reflection SFT on WISE, OneIG-Bench, and T2I-CompBench++, none used in training, and only the unified model runs at test time. Unlike pipelines that pair an external critic with a separate renderer and keep the critic online at inference (Idea2Img, ReflectionFlow, SLD, GenArtist), both roles here are one policy and the verifier is training-only. WISE +10.97, CompBench +4.63, OneIG +3.48; an external GPT-5.5 critic pipeline reaches 79 on GenEval with a 27.6% repair rate, against 84 and 64.9% for UMM-Reflection.
Mechanistic analysis indicates RL does not create a new capability but steers the policy toward a region the backbone already has: a held-out linear probe for correctness on the understanding stream reaches AUC 0.804/0.807/0.815 under Base/SFT/RL, while the share of initially failing images inside that readout's dense pass region rises from 34% to 62% (SFT: 36% to 42%). It supplies representation-level evidence for the 'RL selects rather than creates' reading and helps explain why 1,000 updates on a 3,000-prompt pool suffice. All 553 trajectories were replayed; on the 153 prompts where both policies used exactly three edits, DINO distance is 0.11 versus 0.28 (paired difference 0.17, 95% CI [0.13, 0.21]); with image, renderer, and noise fixed and only the instruction swapped, the RL instruction repairs 48.4%, the original request 20.5%, and a rule-written instruction 27.6%.
Perspective
The result targets unified multimodal models that both understand and generate, especially backbones with flow-matching renderers; training needs a frozen verifier to supply graded rewards, while inference needs only the model itself. It is validated on BAGEL with a 3,000-prompt GenEval-style pool, 1,000 updates, at most three repair rounds, training resolution 512 and 50 denoising steps at inference, for compositional text-to-image alignment repair. Teams aiming to reduce reliance on external critics and fold multi-round reflection into the training loop can reuse the protocol and reward design directly; counting edits are flagged by the authors as the next training target, to which the same reward and protocol apply unchanged.
Evaluation uses a controlled protocol (instruction-voice prompts, 512 resolution, one image per prompt) that differs from native leaderboard protocols, so cross-paper comparisons need care; OneIG reports question-dependent alignment on 695 eligible prompts rather than an overall score, and WISE reports a weighted group aggregate rather than pooled accuracy. The counting family shows flat training reward across 1,000 updates and unchanged GenEval accuracy at 67.5, indicating the current edit actions do not yet cover it. In addition, SFT trajectories are distilled from external models (GPT-5.5 and Qwen-Image series), so the style and coverage of the supervision shape the cold start; stopping behavior is also sensitive, since adding a correct-DONE bonus makes the policy stop prematurely and lowers accuracy from 82 to 79.
