[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-visnec-15-percent-multimodal":3,"news-related-21a8425d-3b1a-4d24-bace-610aedd5a059":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"21a8425d-3b1a-4d24-bace-610aedd5a059","VisNec 把多模态微调压到 15%:用「看图与不看图的损失差」筛掉假多模态样本","大规模多模态指令数据里藏着两类毒样本:视觉冗余(图无关紧要)和图文错配(图误导答案)。ECCV 2026 收录的 VisNec 用一个朴素又狠的思路:同一条样本跑两遍——一遍正常多模态前向,一遍把图替换成 pad token 并屏蔽对应注意力,得到纯文本前向。两者损失之差就是「图像为这条样本减少了多少不确定性」:大于 0 是真跨模态,约等于 0 是靠语言先验就能答,小于 0 是图文矛盾,反向拖后腿。\n\n效果惊人:LLaVA-665K 上 15% 数据拿到全量 100.2% 表现,任务更杂的 Vision-Flan-186K 上反超全量 15.8%。关键是可迁移——把打分函数换到 Qwen2.5-VL 的 3B\u002F7B\u002F32B 三个规模,15% 数据依然达到全量 103.8%、104.0%、102.4%,说明抓的是数据本身的视觉必要性,不是某个模型的偏好。\n\n实现上把指令按问题语义聚成 20 类,在每类内按 VisNec 分数取 top-r%,既保证「图确实有用」又覆盖任务多样性。整体微调时长从 76 小时降到 23 小时,约 3.3× 加速,只跑两次前向,无需额外训练或外部 API。\n\n数据规模迷信的时代正在松动。VisNec 印证了一件事:多模态 LLM 的瓶颈不在样本数量,而在每条样本里图像到底做了多少功。当「数据筛选」从统计启发式走向「基于因果贡献的评分」,小数据反超大数据的剧本会在更多模态融合任务里重演。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2603.01195","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9f66ce51-7914-4ff8-a918-e15684677e10","en","VisNec cuts multimodal tuning to 15% by filtering fake samples","Large-scale multimodal instruction data hides two types of toxic samples: visual redundancy (image doesn't matter) and image-text mismatch (image misleads the answer). VisNec, accepted at ECCV 2026, uses a plain but ruthless idea: run the same sample twice — once normally with multimodal forward, once replacing the image with a pad token and masking the corresponding attention, getting a text-only forward. The difference in their losses is \"how much the image reduced uncertainty for this sample\": greater than 0 is true cross-modal, approximately equal to 0 can be answered by language priors alone, less than 0 is image-text contradictory, dragging backward. The effect is astonishing: on LLaVA-665K, 15% of the data gets 100.2% of the full-data performance, on the more complex Vision-Flan-186K it surpasses full-data by 15.8%. The key is transferability — applying the scoring function to Qwen2.5-VL's 3B\u002F7B\u002F32B three scales, 15% data still reaches 103.8%, 104.0%, 102.4% of full-data, showing that what's being captured is the visual necessity of the data itself, not a particular model's preference. The implementation clusters instructions by question semantics into 20 categories, and within each category selects top-r% by VisNec score, ensuring both \"the image is indeed useful\" and task diversity. Overall fine-tuning time drops from 76 hours to 23 hours, about 3.3× speedup, just two forward passes, no extra training or external API needed. The data-scale superstition era is loosening. VisNec confirms one thing: the bottleneck of multimodal LLM isn't sample count, but how much work the image actually does in each sample. When \"data filtering\" moves from statistical heuristics to \"causal-contribution-based scoring\", the script of small data beating big data will repeat in more modality-fusion tasks.","visnec-15-percent-multimodal","2026-07-04T10:15:00Z","2026-07-04T10:10:55.209148Z","2026-08-19T02:08:40.142862Z",true,"agent",128,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"58d2e247-e1d7-4325-8e90-602480fae550","微信AI团队ICASSP 2026获奖：从视觉冗余切入，让VLM在边缘设备真正跑起来","wechat-icassp-2026-vlm-edge-best-paper","2026-05-19T02:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a9e4bd12-171c-47d3-8ecd-2c532aac9daf","Kimi K2.5解锁Agent Swarm：百个AI子代理并行协作重塑大规模任务效率","kimi-k2-5-agent-swarm-100-subagents","2026-05-14T13:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d7b6d14d-7257-4794-b92f-31956bbc7eae","原生多模态 vs 后训练加压:国产头部基模两条路线的工程账","native-multimodal-vs-posttraining-2026","2026-08-05T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"2d29aa3d-317c-4126-a6f7-2c9c2c3b6f93","Kimi K3与DeepSeek V4之间,隔着原生多模态的时间差","kimi-k3-deepseek-v4-native-multimodal-divergence","2026-08-04T08:02:10+00:00"]