[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-moonshot-perceptionbench-atomic-perception":3,"news-related-ff0bc92a-295a-4707-be8d-76115fe9eeee":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","Moonshot AI 开源原子视觉感知基准 PerceptionBench:从 42 个基准的失败点反推出 10 类能力、3000 道题。16 个前沿多模态模型无一过 60%,总分第一的 GPT-5.6-Sol 幻觉项仅 26.9%,Kimi K3 以 58.5 排第二。","当一个多模态模型把图里的三个杯子数成两个,你很难分清它是\"没看见\",还是\"看见了但没算明白\"。现有视觉基准几乎都答不上这个问题——它们考的是端到端的最终答案,把感知错误和推理错误混在同一张成绩单里。Moonshot AI 这次干脆把两者拆开,发布了一个专门测\"原子视觉感知\"的基准 PerceptionBench,数据集放在 Hugging Face(CC-BY-NC-4.0),评测代码以 Apache 2.0 开源在 [GitHub](https:\u002F\u002Fgithub.com\u002FMoonshotAI\u002FPerceptionBench),方法论见 arXiv:2607.24957。\n\n## 从 42 个基准的失败点反推\n\nPerceptionBench 最有意思的不是分数,而是造题方法。团队没有先验地定义\"视觉感知应该考什么\",而是去回溯前沿模型在 42 个现有基准上最早出错的那个点,把失败归因到感知层面,再从归因结果里蒸馏出十类原子能力:视觉关系、计数、属性、深度与 3D、定位、比较、细粒度识别、上下文整合、OCR,以及感知幻觉。每道题只考一种能力,答案短且唯一,难度来自\"看\"本身而不是推理或知识。最终发布的 3000 道验证题,是从 17000 多个内部验证样本里按能力均衡和难度分层抽出来的;其中 60% 由源基准的失败案例分解而来,40% 围绕补充图片新写。\n\n## 16 个前沿模型,没有一个过 60%\n\n结果相当难看。16 个前沿多模态模型(10 个闭源、6 个开源)统一提示词、放开最大推理预算跑完全部题目,没有一个总体准确率过 60%:GPT-5.6-Sol 以 59.7 排第一,Kimi K3 58.5 第二,Claude-Fable-5 57.2 第三,Gemini-3.1-Pro 56.2,GPT-5.5 55.8。开源侧表现最好的 Qwen3.7-Plus 是 51.1;榜单下半区,Grok-4.5 只有 41.0,GLM-5V-Turbo 39.6,Minimax-M3 33.1,垫底的 GLM-4.6V 32.5。评分环节用 GPT-oss-120B 做裁判,在 300 样本人工审计下与人类判断的一致率为 99.7%。\n\n更有信息量的是分项。总分接近的模型,能力剖面可能完全不同:GPT-5.6-Sol 的定位项拿到 76.7,Gemini-3.1-Pro 同项只有 52.7,两者总分却只差 3.5 分。总榜第一的 GPT-5.6-Sol 在感知幻觉这一项上只有 26.9——它经常\"看见\"图里不存在的东西;而总分只有 52.0 的 Gemini-3.5-Flash,幻觉项反而有 50.6。单看总榜,这些差异全部不可见。\n\n## 猜对,不等于看见\n\n论文还有一个不太显眼但很关键的观察:相当一部分答对的题,换个问法再问一遍就守不住了。这说明模型很多时候在模式匹配而不是真的感知——凭语言先验猜出\"图里大概有什么\",而不是把像素读出来。对依赖 VQA、MMBench 这类整体性基准选型的人,这是个提醒:语言能力强的模型可以用先验补偿感知短板,把整体分\"考\"上去。此外还有个结构性问题:各源基准捕捉到的感知错误切片重叠很弱,两两加权 Jaccard 均值只有 0.20——没有任何一组现有基准能近似覆盖感知全貌,想靠\"多拼几个现有评测\"凑出感知维度,拼不出来。\n\n## 所以呢\n\n对做多模态系统的团队,PerceptionBench 的正确用法是当诊断工具:一个总分 55 但深度感知只有 30 出头的模型,会直接告诉你训练数据该往哪补、架构该改哪。对其他人,这份榜单至少澄清了一件事——当下前沿模型的瓶颈未必在推理,而在更上游的\"看\"。连总分第一的模型都会把不存在的东西看成存在,你敢放心让它读片、验工件、盯监控吗?","https:\u002F\u002Fwww.kimi.com\u002Fblog\u002Fperception-bench","0ec8f614-42c7-4256-8591-209e1e39eb6b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5226bda8-2fb0-4a26-a89e-ddc2d0b3d8e7","en","PerceptionBench: no frontier multimodal model clears 60% on visual perception","Moonshot AI open-sources PerceptionBench, an atomic visual perception benchmark reverse-engineered from failure points across 42 existing benchmarks into 10 capabilities and 3,000 questions. None of 16 frontier multimodal models clears 60%; overall leader GPT-5.6-Sol scores just 26.9 on hallucination, with Kimi K3 second at 58.5.","When a multimodal model counts three cups in an image as two, you can't easily tell whether it failed to *see* or failed to *reason*. Almost no existing visual benchmark answers that question — they test end-to-end answers and mix perceptual errors with reasoning errors on the same scoreboard. Moonshot AI has now split the two apart with PerceptionBench, a benchmark built specifically for \"atomic visual perception.\" The dataset is on Hugging Face (CC-BY-NC-4.0), the evaluation code is open-sourced under Apache 2.0 on [GitHub](https:\u002F\u002Fgithub.com\u002FMoonshotAI\u002FPerceptionBench), and the methodology is described in arXiv:2607.24957.\n\n## Built From 42 Benchmarks' Failure Points\n\nThe most interesting thing about PerceptionBench isn't the scores — it's how the questions were made. Instead of pre-defining what \"visual perception\" should test, the team traced frontier models' earliest failure points across 42 existing benchmarks, attributed those failures to perception, and distilled the attributions into ten atomic capabilities: visual relation, counting, attribute, depth & 3D, localization, comparison, fine-grained recognition, context integration, OCR, and perception-related hallucination. Each question tests exactly one capability, with short, uniquely determined answers — difficulty comes from seeing, not from reasoning or knowledge. The released set of 3,000 verified questions was subsampled with capability-level balancing and difficulty stratification from an in-house pool of 17,000+ verified samples; 60% are atomic sub-questions decomposed from attributed failures on source benchmarks, and the remaining 40% were newly authored on supplemented images.\n\n## Sixteen Frontier Models, None Clears 60%\n\nThe results are rough. Sixteen frontier multimodal models (ten proprietary, six open-source) ran the full set with unified prompts and the highest available reasoning budget — and none reached 60% overall accuracy. GPT-5.6-Sol leads at 59.7, Kimi K3 takes second at 58.5, Claude-Fable-5 third at 57.2, Gemini-3.1-Pro 56.2, GPT-5.5 55.8. The best open-weight performer is Qwen3.7-Plus at 51.1; in the lower half, Grok-4.5 manages only 41.0, GLM-5V-Turbo 39.6, Minimax-M3 33.1, with GLM-4.6V at the bottom on 32.5. Scoring used GPT-oss-120B as judge, with 99.7% agreement with human judgment on a 300-sample audit.\n\nThe subscores carry the real information. Models with nearly identical overall scores can have completely different capability profiles: GPT-5.6-Sol scores 76.7 on localization while Gemini-3.1-Pro gets 52.7 on the same category, yet their overall scores differ by just 3.5 points. The overall leader GPT-5.6-Sol scores only 26.9 on perception-related hallucination — it frequently \"sees\" things that aren't there — while Gemini-3.5-Flash, sitting at 52.0 overall, scores 50.6 on hallucination. None of this is visible on a single leaderboard.\n\n## Guessing Right Isn't Seeing\n\nThe paper also contains a quiet but important observation: a large share of correctly answered questions don't survive being asked again. Models are often pattern-matching rather than perceiving — inferring what an image probably contains from language priors instead of reading the pixels. For anyone selecting models on holistic benchmarks like VQA or MMBench, this is a warning: models with strong language ability can compensate for weak perception with priors and \"test\" their way to a higher overall score. There's also a structural problem: the perception-error slices captured by existing benchmarks overlap only weakly, with a mean pairwise weighted Jaccard of 0.20 — no single benchmark, or small group of them, approximates perception as a whole. You can't assemble a perception axis by stapling together a few existing evals.\n\n## So What\n\nFor teams building multimodal systems, PerceptionBench works best as a diagnostic: a model scoring 55 overall but ~30 on depth perception tells you exactly where to invest in training data or architecture. For everyone else, the leaderboard clarifies one thing — the bottleneck of frontier models may not be reasoning, but the \"seeing\" upstream of it. If even the top-scoring model sees things that don't exist, would you trust it with radiology reads, factory inspection, or surveillance monitoring?","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00Z","2026-08-26T13:13:47.267659Z","2026-08-26T13:13:47.267670Z",true,"agent",21,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"5a274662-0c3f-492c-a0e8-a46c5a0be783","别再自己给自己打分了:Co-RL 让模型互相判卷,无标签 RL 追平有监督","co-rl-peer-reward-label-free-rl","2026-08-20T19:10:00+00:00"]