TRM first writes a case-adaptive rubric before scoring, beating open-source reward models and nearing proprietary ones on image generation and editing benchmarks
Synopsis
The work introduces the "Think Before You Score" paradigm and the Thinking Reward Model (TRM), which first generates a case-adaptive rubric for each task condition and candidate output, then inspects the candidate criterion by criterion and aggregates the evidence into a fine-grained pointwise reward; it also proposes PD-GRPO to use pairwise preference supervision for better discrimination while mitigating score polarization, and reports state-of-the-art results among open-source reward models on image generation and editing reward-modeling benchmarks, highly competitive with proprietary alternatives, with TRM-guided reinforcement learning consistently improving diverse visual generation models.
Interpretation
TRM reframes visual reward modeling from mapping conditions and candidates directly to a scalar score into first determining what should be evaluated for this case and then judging how well the candidate performs, structured as determine-inspect-aggregate-score with case-adaptive rubrics, rubric-level judgments, dimension-level assessments, and a final pointwise reward. Existing visual reward models largely rely on fixed evaluation criteria, and even reasoning-based scoring rarely makes explicit what should be checked for a particular case; TRM instantiates rubrics from the task condition and candidate output while keeping a shared high-level structure (Prompt Alignment and Visual Quality for both tasks, plus Aesthetics for generation and Source Consistency for editing). The paper uses a comparison table to show that prior methods generally lack adaptive rubrics, and defines TRM's process and three-dimension instantiation; this is a paradigm-and-structure argument supported by the later benchmark results.
The authors build roughly 20K image generation cases and 28K image editing cases (about 48K total) through a unified pipeline and a two-stage human-in-the-loop annotation process, yielding structured supervision with case-adaptive rubrics, criterion-level judgments, and quality scores. Unlike data that provide only scalar preferences or scores, this provides learnable structured evaluation traces covering multiple tasks, difficulty levels, model sources, and failure modes, with editing data rebalanced by original ratings to emphasize intermediate quality levels. Data scale, source benchmarks (such as Edit-Compass, UniREdit-Bench, Qwen-Image-Bench, EvalMuse), and the two-stage annotation pipeline are described in the main text and appendix; quantitative annotator-agreement metrics are not given in the text.
Observing that a Bradley-Terry-style objective keeps enlarging score margins even after the preference ordering is correct, which can induce score polarization, the authors propose PD-GRPO: for each preference pair, pointwise evaluations are sampled independently for both candidates and each rollout is credited against the mean score of the opposite group, so no further benefit accrues once the required separation margin is met. The BT-style objective has no interior optimum in score space and drives the two sides toward opposite boundaries; PD-GRPO's reward becomes constant once the margin is satisfied, preserving the pointwise inference interface while exploiting pairwise supervision for fine-grained discrimination, with the margin adjustable by pair difficulty. The paper derives and illustrates the monotonicity and polarization of BT, and specifies PD-GRPO's reward construction and group-relative advantage computation; the separation margin is set to 0.5 for image editing and 1.0 for image generation on the normalized score scale.
Experiments show TRM reaches state-of-the-art performance among open-source reward models and remains highly competitive with proprietary alternatives on image generation and editing reward-modeling benchmarks, and that as a FlowGRPO reward it consistently improves BAGEL, FLUX.1-dev, SD3.5-M, and editing models including Flux2-Klein 4B/9B, Flux-Kontext, and SenseNova-U1.5. On generation, TRM(RL) reaches 71.2%/67.9% on GenAI-T2I/MMRB2-T2I versus 70.1%/65.8% for SFT and 58.9%/59.4% for the baseline; on editing, TRM(RL) reaches 0.786/0.674/0.773 on EditScore-ERB, 58.2% on MMRB2, 71.3% on EditReward-ERB, and 0.640/0.660 on EditReward-Compass, with a 9B model outperforming the 72B EditScore; in downstream RL, BAGEL gains 5.81 points on TIIF-Long and FLUX.1-dev gains 6.76 and 6.56 points on TIIF-Short and TIIF-Long. Results span multiple public benchmarks and baseline families (GPT, Gemini, Qwen series, plus HPSv3, UnifiedReward, RationalRewards, EditScore, FIRM-Reward), with ablations (same BAGEL backbone against AlphaGRPO, same training setting against EditScore-72B) and qualitative examples; the main generation table reports accuracy over non-tied predictions, with tie-aware results in the appendix.
Perspective
The work targets reward modeling and RL post-training for two visual generation tasks: image generation, where the condition is a text prompt, and image editing, where the condition is a source image plus an editing instruction, with a shared evaluation structure of Prompt Alignment and Visual Quality plus a task-specific dimension. It suits settings that need fine-grained, interpretable pointwise rewards to drive FlowGRPO-style online optimization, and teams that want a smaller reward model to stand in for large specialized ones; the paper also describes asynchronous reward computation and shared-prefix caching, indicating that deployment inside real RL training is part of the intended setting.
The text is a parse of the full paper, but some values appear as placeholders in the main text (for example specific learning rates, KL coefficients, and some margin and batch settings), so exact hyperparameters still need to be confirmed in the original appendix. In addition, the main generation accuracy excludes predicted ties while tie-aware results are given in the appendix, so readers may want to compare conclusions under both protocols; the degree of expert calibration agreement in the annotation pipeline, and how case-adaptive rubrics behave on tasks outside the training distribution, are not quantified in the text and remain open questions to watch.
