ByteDance Seed × University of Hong Kong × Peking University propose SpectraReward in arXiv 2607.11886, using "reading back the prompt" as a zero-shot reward for T2I RL; Self-SpectraReward lets BAGEL reward itself, with GenEval 89.5 beating Qwen3-VL-235B-A22B; the core conclusion: reward-policy alignment matters more than reward scale.