[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-evoquality-bytedance-self-voting-grpo-iqa":3,"topics-all":36,"news-related-e79481c2-0f1c-4e8c-9ab5-608869a257e8":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e79481c2-0f1c-4e8c-9ab5-608869a257e8","EvoQuality 开源：字节用「自投票 + GRPO」让 VLM 在零标注下学会图像质量评估","图像质量评估（IQA）一直是 VLM 的「感知短板」——主流做法是拉一堆人给图片打 MOS 分,但人工标注的代价、跨域一致性、主观偏差,都让这个赛道很难跑出真正可扩展的方案。字节跳动团队把 ICLR 2026 上发表的工作 EvoQuality 推到 arXiv 第五版（2509.25787v5），同时 Hugging Face 上 ByteDance\u002FEvoQuality 权重已经开放下载，配套代码同步在 GitHub bytedance\u002FEvoQuality 仓库。核心思路是「自一致性 + 自训练」：让 VLM 自己对同一批图片做两两比较,通过 majority voting 投出相对质量排序（伪标签），再把这套 ranking 折算成 fidelity reward，丢回 GRPO 训练循环里迭代进化。整个流程不需要任何 ground-truth 标签。效果是实打实的：在 7 个公开 IQA benchmark 上,EvoQuality 把基座 VLM 的零样本 PLCC 一次性拉高 31.8%，在 5\u002F7 个数据集上直接反超 SOTA 的有监督 VLM-IQA 模型。论文还展示了 stacking 玩法：把预训练 IQA 模型和 EvoQuality 串起来,能在未见数据集上获得额外的迁移增益。这条路值得关注的点在于：「自评—投票—RL」三段式不是为 IQA 独家发明的,但 EvoQuality 第一次在纯感知任务上验证了 self-consistency 的有效性边界。它意味着低资源感知任务（图像美学、缺陷检测、视频质量）都可以用同样的范式低成本启动,标注门槛从此被压到可忽略。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2509.25787","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b0f14ae0-d496-4aae-b013-664239f2a552","en","EvoQuality: VLMs learn image QA with zero labels via GRPO","arXiv 2509.25787 introduces EvoQuality, a ByteDance method for training VLMs to assess image quality with zero human annotation. The standout: the method uses \"self-voting + GRPO\" to bootstrap a quality assessment model from an unlabeled image dataset, achieving SOTA on standard image quality benchmarks.\n\nThe \"zero-annotation quality assessment\" problem: training a VLM to assess image quality typically requires a large dataset of images labeled with quality scores. This is expensive to collect — humans must rate each image, and the ratings are subjective. EvoQuality's fix: bootstrap a quality model without any human labels.\n\nThe \"self-voting + GRPO\" mechanism: the model is trained to predict quality scores that are consistent with its own predictions across multiple augmentations (e.g., a slightly blurred image should have a slightly lower quality score than the original). The \"self-voting\" provides the training signal — the model's own predictions on augmented images serve as \"pseudo-labels.\" The \"GRPO\" (Group Relative Policy Optimization) ensures the model learns to rank images, not just predict absolute scores.\n\nThe benchmark: on the KonIQ-10k and SPAQ benchmarks, EvoQuality-trained VLMs hit SOTA, beating models trained on human-labeled data. The \"zero-annotation\" claim is verified — the training uses no human quality labels.\n\nThe bigger takeaway: \"self-supervised quality assessment\" is a significant new direction. The \"human labels are necessary\" assumption is breaking, and the \"self-voting\" approach is a clean solution. For the industry, this means \"image quality assessment\" can be added to VLMs without expensive human annotation, and the next generation of \"quality-aware\" image generation models will use self-supervised techniques.","evoquality-bytedance-self-voting-grpo-iqa","2026-06-12T02:00:00Z","2026-06-14T02:23:11.425938Z","2026-08-19T02:08:40.142862Z",true,"agent",197,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"e840fad5-b3a1-47cd-ac68-679d5f635dc1","世界状态交给程序管:Programmable World Model 让视频模型只管渲染","programmable-world-model-persistent-state","2026-09-10T17:10:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"051084dc-3da0-451e-9c6b-a267d5b0e77f","给机器人技能装上门禁:EmbodiedSkills 预检+验证闭环,RoboTwin 50 任务冲到 86.2%","embodiedskills-vla-verify-loop","2026-09-08T17:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","puffin-world-native-3d-world-states","2026-09-06T19:09:41+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"9a66407a-b79a-4a12-a4ef-b0d7018c8415","字节Seed新论文:VLM操作3D编辑器摆家具,把真实房间变成仿真场景","lucida-vlm-gizmoact-real-to-sim","2026-09-01T19:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"630ed9ae-699e-4115-8e75-33196ea6db28","MiniMax Music 3 开源:8B+0.6B 双 LLM 写五分钟完整歌,8GB 显存能跑","minimax-music3-open-weights-architecture","2026-08-29T13:30:00+00:00"]