[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ruler-rubric-rewards-svg-generation":3,"topics-all":38,"news-related-aa9b279e-07c6-4d0c-86f0-403c4af321fa":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"aa9b279e-07c6-4d0c-86f0-403c4af321fa","蚂蚁 RULER:六维评分表给 SVG 生成重写奖励信号","SVG 生成没有标准答案,CLIP 等标量指标用在矢量内容上会失真,当 RL 奖励还会诱发作弊。蚂蚁团队提出 RULER:每条指令转成六项评分细则,判分模型逐项打分加权成 GRPO 奖励,不需要配对真值和人工偏好。MMSVG 两基准评分从 0.432\u002F0.395 升到 0.693\u002F0.683。","让语言模型写 SVG 代码这件事,卡在一个很根本的地方:它没有标准答案。一张\"画个带睫毛的眼睛图标\"的作业,写成代码可以有无数种都对的做法,既没有配对的参考图,也没有人工偏好标签。于是评估和训练都只能借用在自然图像上校准的标量指标——CLIP 分、美学分——但矢量内容风格化程度高,这些指标迁移过去就会失真。论文给了一个直观例子:一张忠实的蓝色相机图标和一张明显损坏的 SVG,CLIP 给出 0.18 和 0.23(坏的反而更高),美学分 4.93 对 4.84 也几乎无区分度,而人工设计的评分细则能给到 0.94 对 0.30,方向正确。更麻烦的是,把这类失真指标直接当强化学习奖励,会触发 reward hacking:模型学会讨好打分器,而不是把图做好。\n\n## 蚂蚁的解法:把\"标准答案\"换成\"评分细则\"\n\n蚂蚁集团联合港科大(广州)、牛津等机构的团队提出 RULER,思路是给每条生成指令动态生成一份六项评分细则(rubric),覆盖语义、视觉、渲染风格三个维度,每个维度两项:语义保真看概念可读性、主要部件、提示特有关系;视觉质量看轮廓形态与构图;渲染风格看完成度与风格一致性。细则由前沿模型(Claude-Opus-4.6)从纯文本指令推导,训练时用判分 VLM(Qwen3-VL-8B)对渲染后的输出逐项打分,加权求和形成细粒度奖励,再用 GRPO 优化策略模型(Qwen3-8B)。整条管线不依赖任何 SVG 真值参考,也不需要人工标注偏好。\n\n## 结果:小模型追平大它数倍的 DeepSeek-V3\n\n在 MMSVG-Illustration 和 MMSVG-Icon 两个基准上,RULER 把评分细则得分从 0.432\u002F0.395 提升到 0.693\u002F0.683,超过各专用 SVG 模型,并追平了大得多的 DeepSeek-V3。盲测人类偏好上,150 个 prompt 的非平局胜率全部过半:对 Qwen3-8B 基线 89.7%,对 Qwen3-32B 66.1%,对 VectorFusion 53.3%,对 OmniSVG 66.9%,对 JanusCoder 96.5%。判分器与人类判断的对齐度也拉得开:Spearman 相关 0.7929(美学分 0.6051、CLIP 0.5518),Goodman-Kruskal γ 成对排序一致率 0.7574。消融实验里去掉任何一个维度的细则都会掉分,验证六项设计缺一不可;换用 GPT-5.5 生成细则、或在 4B\u002F8B 两种底座上跑,结论保持稳健。\n\n## 所以呢\n\n这项工作真正值得记的点,不在 SVG 本身,而在\"开放式生成任务的 RL 奖励该怎么造\":当任务没有 ground truth、标量指标又会失真时,用 LLM 生成的实例感知评分细则当奖励,是一条无需人工标注的可行路径。对做图像生成、UI 代码生成、创意写作这类\"好答案不唯一\"任务的团队来说,这套\"指令→细则→逐项判分→GRPO\"的配方可以直接迁移。需要留意的是,细则质量和判分 VLM 的能力成为新的隐性上限,论文的消融也显示细则设计是 RL 的活性杠杆——打分器有多可靠,奖励就有多可靠。\n\n参考:[arXiv:2609.25270](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25270) · 项目页 hangyuran.github.io\u002FRULER","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25270","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"7e37f23d-605b-4569-9422-cf2fd22cfb2f","en","Ant Group RULER: Rubric Rewards for SVG Generation","Ant Group's RULER replaces scalar SVG rewards with six-item rubrics. MMSVG scores rose from 0.432\u002F0.395 to 0.693\u002F0.683, matching DeepSeek-V3.","Getting language models to write SVG code is stuck at a fundamental point: the task has no ground truth. An assignment like \"draw an eye icon with eyelashes\" admits countless valid code solutions, with no paired reference images and no human preference labels. So both evaluation and training borrow scalar metrics calibrated on natural images — CLIP scores, aesthetic scores — but stylized vector content distorts them. The paper offers a vivid example: a faithful blue camera icon versus a visibly broken SVG get CLIP scores of 0.18 and 0.23 (the broken one scores higher), and aesthetic scores of 4.93 vs 4.84 are nearly indistinguishable, while a hand-designed rubric correctly assigns 0.94 vs 0.30. Worse, plugging such distorted metrics directly in as reinforcement learning rewards triggers reward hacking: the model learns to please the scorer instead of making the image good.\n\n## Ant Group's Fix: Swap Ground Truth for a Scoring Rubric\n\nA team from Ant Group with HKUST (Guangzhou) and Oxford researchers proposes RULER. The idea: dynamically generate a six-item scoring rubric for each generation instruction, spanning semantic, visual, and rendering-style dimensions — two items each. Semantic fidelity covers concept readability, major components, and prompt-specific relations; visual quality covers silhouette, form, and composition; rendering style covers finish and stylistic coherence. The rubric is derived from the text instruction alone by a frontier model (Claude-Opus-4.6); during training a judge VLM (Qwen3-VL-8B) scores rendered outputs item by item, and the weighted sum forms a fine-grained reward optimized via GRPO over the policy model (Qwen3-8B). The whole pipeline depends on no SVG ground-truth references and no human preference annotation.\n\n## Results: A Small Model Matches the Much Larger DeepSeek-V3\n\nOn MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432\u002F0.395 to 0.693\u002F0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3. In blinded human preference over 150 prompts, non-tie win rates clear 50% against every baseline: 89.7% over the Qwen3-8B baseline, 66.1% over Qwen3-32B, 53.3% over VectorFusion, 66.9% over OmniSVG, and 96.5% over JanusCoder. Judge alignment with humans also opens a wide gap: Spearman correlation of 0.7929 (vs 0.6051 aesthetic, 0.5518 CLIP) and Goodman-Kruskal gamma pairwise ranking agreement of 0.7574. Ablations show that removing any rubric axis drops scores — all six items earn their place — and results hold when swapping in GPT-5.5 as rubric generator or running 4B\u002F8B base models.\n\n## So What\n\nThe point worth remembering is not about SVG per se, but about how to construct RL rewards for open-ended generation tasks: when a task lacks ground truth and scalar metrics distort, LLM-generated instance-aware rubrics offer a human-annotation-free path. Teams working on image generation, UI code generation, or creative writing — anywhere \"good answers are not unique\" — can port this instruction-to-rubric-to-itemized-judging-to-GRPO recipe directly. One caveat: rubric quality and judge-VLM capability become the new implicit ceiling, and the paper's ablations identify rubric design as the active lever for RL — the reward is only as reliable as the scorer.\n\nReference: [arXiv:2609.25270](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25270) · Project page hangyuran.github.io\u002FRULER","ruler-rubric-rewards-svg-generation","2026-09-23T23:06:24Z","2026-09-23T23:09:18.077986Z","2026-09-23T23:09:18.078001Z",true,"agent",262,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4f5c4072-660c-47e7-9d9a-9782eb5591bf","Pistis 报告:IDRL 让蒸馏和 RL 交替上岗","pistis-idrl-interleaved-distillation-rl","2026-09-25T19:11:33+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"ee7c1b35-e8cc-41e3-8b8f-27a512a9f639","TempCloze 视频「完形填空」:31 款 Video-LLM 横评,开源模型时间对齐平均 26.54% 逼近乱猜","tempcloze-video-llm-temporal-alignment","2026-09-13T23:06:41+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00"]