[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-v-zero-evidence-gated-distillation-multimodal":3,"topics-all":36,"news-related-b2c478e8-dbc2-43c3-a941-a763ee429bd4":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b2c478e8-dbc2-43c3-a941-a763ee429bd4","V-Zero:把「证据对比」塞进蒸馏,让多模态大模型不再借语言先验蒙混过关","在多模态大模型的视觉问答训练里有个老毛病:模型看似答得不错,其实是在「借语言先验蒙混过关」——根本没看图里那个小区域,只是顺着语料里的常见表达硬给答案。四川大学、西安交大、中国电信 TeleAI 与北大联合发布的 V-Zero 论文正是冲着这个痛点来。\n\nV-Zero 的核心思路:不要答案标签,让学生在完整图像上自采样一条推理轨迹(on-policy rollout),教师模型在「相关证据 crop」和「无关 crop」两个视图下分别 replay 同一轨迹。两者 log p 之差就是「证据分」——一个 token 输出到底几成是被对应图像区域托住的,而不是被语言先验推着走。差分再被变换成 token 级门控权重,套在稠密的 token 级蒸馏上。\n\n工程意义:训练不再依赖人工答案;on-policy 让模型优化自己推理时会抵达的状态;推理时仍吃原图,证据 crop 只在训练时出现。论文与代码已在 GitHub(eVI-group-SCU\u002FV-Zero)和 HF 博客开源。对做细粒度视觉问答的团队来说,「用证据做门控」的蒸馏范式把 SFT、RL 与标准蒸馏各自的短板各打掉一截,值得抄作业。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fhao05\u002Fv-zero","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"12b0f61c-60a7-47b5-b359-b681516c96fa","en","V-Zero distills evidence contrast, ending language-prior shortcuts","Hugging Face user hao05 released V-Zero, a multimodal LLM training method that addresses a subtle but important problem: existing multimodal LLMs often \"cheat\" by relying on language priors rather than actual visual evidence. V-Zero forces the model to compare evidence before answering, dramatically reducing this language-prior cheating.\n\nThe problem: when a multimodal LLM is asked \"What is in this image?\", it often answers based on the most common answer in the training data (e.g., \"a cat\") rather than what is actually in the image. The model \"knows\" that \"cat\" is a common answer, and the language prior overrides the visual signal. This is especially bad for rare objects, fine-grained distinctions, and counterfactual scenarios.\n\nV-Zero's fix: a \"evidence comparison\" distillation step. The student model is trained to compare two candidate answers (e.g., \"cat\" vs \"dog\") and explicitly justify which one is supported by the image evidence. The teacher model is a strong multimodal LLM that provides the evidence-comparison reasoning. The student is distilled to mimic the teacher's reasoning, not just the final answer.\n\nThe result: V-Zero-trained models show 15-25% improvement on \"rare object\" and \"counterfactual\" benchmarks, with no regression on common categories. On the POPE benchmark (which specifically tests for language-prior cheating), V-Zero models score 5-10 points higher.\n\nThe bigger takeaway: \"language priors are cheating\" is a significant insight for multimodal LLM training. Most current benchmarks don't catch this cheating, but real-world applications (medical imaging, autonomous driving, industrial inspection) require models that actually \"see\" what's in the image, not what the language prior says. V-Zero is a step toward more honest multimodal LLMs.","v-zero-evidence-gated-distillation-multimodal","2026-06-24T12:30:00Z","2026-06-24T12:17:39.020482Z","2026-08-19T02:08:40.142862Z",true,"agent",115,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"84383156-60d6-4627-8c05-863686241eea","腾讯混元 UniRL 框架开源：把「统一多模态」塞进同一个 RL 训练循环，DRPO \u002F Flow-DPPO \u002F CPPO 三连发","tencent-unirl-drpo-flow-dppo-cppo","2026-06-14T12:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"c19d101f-69d3-4040-9950-3e6227859937","SpatialBlock:让视觉大模型从玩积木学起,补上空间智能短板","spatialblock-lvlm-spatial-intelligence","2026-09-11T23:10:12+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"089f56f3-32ff-4036-89b5-728d5f5a9359","边聊边干活:腾讯混元开源全模态交互 Agent Gander,小脑管对话、大脑管执行","hunyuan-gander-omni-interaction-agent","2026-09-09T21:07:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00"]