RULER uses per-instruction six-item rubrics as RL rewards, lifting Qwen3-8B SVG generation from 0.432 to 0.693 and matching the much larger DeepSeek-V3
Synopsis
The work first shows, on 900 human-rated SVG samples generated by Claude-Opus-4.6, Qwen3-32B, and Qwen3-8B, that prompting a vision-language judge with a multi-axis rubric correlates with human judgment far better than scalar metrics such as CLIP and aesthetic scores (Spearman 0.7929, Goodman-Kruskal Gamma 0.7574), then introduces RULER: it derives a six-item instance-aware rubric from the text instruction alone, has a judge VLM score each rendered rollout item-by-item, and optimizes the weighted reward with GRPO, raising the rubric score on MMSVG-Illustration and MMSVG-Icon from the Qwen3-8B backbone's 0.432 and 0.395 to 0.693 and 0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3 without paired SVG ground truth or human preference labels.
Interpretation
The work replaces single-scalar quality assessment for open-ended SVG generation with multi-axis rubric scoring and quantifies its alignment with human judgment: over 900 rendered samples, rubric scoring reaches a sample-level Spearman correlation of 0.7929 and a pairwise ranking agreement (Goodman-Kruskal Gamma) of 0.7574, both above aesthetic scores and CLIP. Prior SVG generation work commonly used scalar metrics such as CLIPScore, aesthetic classifiers, and HPS, which were calibrated on natural images and transfer poorly to stylized vector content; this work systematically compares rubric scoring against those scalar metrics on the same sample set. 900 rendered SVG samples, 300 each from three models of varying capability (Claude-Opus-4.6, Qwen3-32B, Qwen3-8B), rated by ten human annotators on a 0-100 holistic quality score with three validators rescoring flagged cases; the alignment analysis uses GPT-5-mini with a shared universal rubric, separate from the Qwen3-VL-8B judge used during training.
RULER converts each instruction into a six-item instance-aware rubric spanning semantic fidelity, visual quality, and rendering style; a judge VLM scores each rendered rollout item-by-item, and the normalized weighted satisfactions form a dense reward optimized via GRPO group-relative advantages. Earlier universal-rubric scoring applies the same query-agnostic checklist to every sample and cannot reflect instruction-specific notions of correctness; RULER's rubric is generated from the text instruction alone, so it needs neither paired SVG ground truth nor human preference labels while preserving the open-ended solution space. The rubric-generation prompt explicitly requires the rubric to remain open-ended and prohibits exact reconstruction requirements on geometry, placement, colors, part counts, or other implementation details; training uses 30,737 prompt-rubric pairs (19,531 for MMSVG-Icon, 11,206 for MMSVG-Illustration) after format and quality filtering that removes rubrics whose ideal SVG scores below 0.9.
On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from the Qwen3-8B backbone's 0.432 and 0.395 to 0.693 and 0.683, surpassing dedicated SVG specialists (e.g., OmniSVG-8B at 0.288 and 0.565) and matching the substantially larger DeepSeek-V3 (0.686 and 0.673), while remaining competitive on CLIP, aesthetic, and HPS metrics. The result indicates that closing the open-ended quality gap requires more than scaling SVG-specific pretraining; an 8B model trained with instance-aware rubric rewards can reach visual quality on par with a much larger model. The comparison table covers three baseline families (diffusion-optimized methods, foundation LLMs, and SVG specialists) on two benchmarks; a blinded pairwise human preference study on 150 MMSVG-Bench prompts shows RULER's non-tie win rate above 50% against all five evaluated baselines, ranging from 53.3% against VectorFusion to 96.5% against JanusCoder.
Reward-design ablations identify rubric design itself as the active lever for RL on open-ended SVG generation: removing the visual-quality axis causes the largest decline, removing the rendering axis the second largest, and removing the semantic axis the smallest; replacing the rubric with a stricter, style-free Rubric-S plus a text-hint penalty drops the rubric score to 0.536 and triggers text-hint hacking where readable prompt text is embedded in the SVG. The work isolates the reward signal by fixing the Qwen3-8B backbone and the GRPO optimizer and varying only the reward across zero-shot, a CLIP+aesthetic+HPS scalar reward, a universal-rubric reward, and the full instance-aware rubric. The scalar multi-metric reward C+A+H RL inflates the aesthetic score to 6.697 (Illustration) and 6.210 (Icon) but degrades CLIP and collapses the Icon rubric score to 0.262, below zero-shot, while blowing up sequence length to 6.3k tokens, a characteristic reward-hacking failure mode; universal-rubric RL improves steadily but leaves rubric gaps relative to the full method on both benchmarks.
Perspective
The result targets open-ended text-to-SVG code generation, for instruction sets that have neither paired SVG ground truth nor human preference labels, with training data drawn from the training splits of MMSVG-Icon and MMSVG-Illustration and evaluation on MMSVG-Illustration, MMSVG-Icon, and MMSVG-Bench. The approach transfers directly to other visual code tasks that require judging rendered output, provided text instructions and a VLM judge that can score rendered results item-by-item are available. Rubric generation is a one-time offline preprocessing step that is cached, and the rubric generator is not queried during policy optimization, so the pipeline suits settings where reward specifications are built in bulk before RL. The authors state they will release the generated rubrics, processed training data, and code, offering a starting point for reproducing or adapting the rubric structure on one's own instruction sets.
Rubrics are generated by a frontier model and scored by a VLM judge, so reward quality may inherit their biases and failure modes; the judge can still overvalue superficial cues or underweight subtle stylistic qualities, especially on prompts outside the training distribution. Each RL step requires rendering sampled SVGs and querying a judge VLM, substantially more expensive than lightweight scalar rewards such as CLIP or heuristic code-based signals, which may limit scaling to larger models, longer rollouts, or broader hyperparameter searches. Decomposing quality into semantic, visual, and stylistic axes works on these benchmarks but may not fully capture all valid artistic intents or domain-specific preferences. Training and test prompts are generated separately yet remain within the same benchmark distribution, and the authors explicitly do not treat this as a domain-shifted out-of-distribution evaluation; rubric generation costs $0.0655 per prompt and $2,013.12 in total, and how that overhead scales to much larger instruction sets remains an open question.
