Getting language models to write SVG code is stuck at a fundamental point: the task has no ground truth. An assignment like "draw an eye icon with eyelashes" admits countless valid code solutions, with no paired reference images and no human preference labels. So both evaluation and training borrow scalar metrics calibrated on natural images — CLIP scores, aesthetic scores — but stylized vector content distorts them. The paper offers a vivid example: a faithful blue camera icon versus a visibly broken SVG get CLIP scores of 0.18 and 0.23 (the broken one scores higher), and aesthetic scores of 4.93 vs 4.84 are nearly indistinguishable, while a hand-designed rubric correctly assigns 0.94 vs 0.30. Worse, plugging such distorted metrics directly in as reinforcement learning rewards triggers reward hacking: the model learns to please the scorer instead of making the image good.

Ant Group's Fix: Swap Ground Truth for a Scoring Rubric

A team from Ant Group with HKUST (Guangzhou) and Oxford researchers proposes RULER. The idea: dynamically generate a six-item scoring rubric for each generation instruction, spanning semantic, visual, and rendering-style dimensions — two items each. Semantic fidelity covers concept readability, major components, and prompt-specific relations; visual quality covers silhouette, form, and composition; rendering style covers finish and stylistic coherence. The rubric is derived from the text instruction alone by a frontier model (Claude-Opus-4.6); during training a judge VLM (Qwen3-VL-8B) scores rendered outputs item by item, and the weighted sum forms a fine-grained reward optimized via GRPO over the policy model (Qwen3-8B). The whole pipeline depends on no SVG ground-truth references and no human preference annotation.

Results: A Small Model Matches the Much Larger DeepSeek-V3

On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3. In blinded human preference over 150 prompts, non-tie win rates clear 50% against every baseline: 89.7% over the Qwen3-8B baseline, 66.1% over Qwen3-32B, 53.3% over VectorFusion, 66.9% over OmniSVG, and 96.5% over JanusCoder. Judge alignment with humans also opens a wide gap: Spearman correlation of 0.7929 (vs 0.6051 aesthetic, 0.5518 CLIP) and Goodman-Kruskal gamma pairwise ranking agreement of 0.7574. Ablations show that removing any rubric axis drops scores — all six items earn their place — and results hold when swapping in GPT-5.5 as rubric generator or running 4B/8B base models.

So What

The point worth remembering is not about SVG per se, but about how to construct RL rewards for open-ended generation tasks: when a task lacks ground truth and scalar metrics distort, LLM-generated instance-aware rubrics offer a human-annotation-free path. Teams working on image generation, UI code generation, or creative writing — anywhere "good answers are not unique" — can port this instruction-to-rubric-to-itemized-judging-to-GRPO recipe directly. One caveat: rubric quality and judge-VLM capability become the new implicit ceiling, and the paper's ablations identify rubric design as the active lever for RL — the reward is only as reliable as the scorer.

Reference: arXiv:2609.25270 · Project page hangyuran.github.io/RULER