[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vlx-seek-1-5":3,"news-related-a36d9d97-42de-4c87-88e9-cdc173b9ab4b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a36d9d97-42de-4c87-88e9-cdc173b9ab4b","VLX-Seek 1.5 把端侧具身感知切成 0.6B\u002F3B\u002F10B 三档：用 None 输出压住目标幻觉","Om AI 联汇 7 月 6 日发布 VLX-Seek 1.5，这是面向端侧具身场景的细粒度感知 VLM 新版本。这次升级的关键词不是「更大」，而是「更可部署」。\n\n新版规划了 0.6B \u002F 3B \u002F 10B 三档模型系列，让无人机、机器狗、监控摄像头等不同算力预算的终端都能选到合适版本——这是把 VLM 一味堆大、最后却塞不进端侧的产品里少见到的工程意识。架构上引入更多 Linear Attention 层和更快的 OPN 候选区域生成，推理时延对端侧更友好。\n\n更值得注意的是「目标幻觉」的处理。具身场景里，机器人错误地追踪一个不存在的目标，代价远超漏检。VLX-Seek 1.5 引入显式 None 输出格式：用户问「图中的 A 和 B」，若 B 不存在，模型必须输出 A 的坐标 + B 的 None。在 HumanRef、VisDrone、RefDrone 三个基准上，Object Hallucination 指标（FP \u002F GT 数量）比上一版和 LocateAnything 都更低。\n\n视觉能力上，新版训练数据加入更多无人机、监控、机器人视角，辅助视觉塔也升级。在 COCO、LVIS、RefCOCO、VisDrone、RefDrone、EmbSpatialBench 等基准上，VLX-Seek 1.5-3B 反超了多个更大的开源\u002F闭源 VLM。\n\nOm AI 联汇宣布开源 10B 版本。在具身感知赛道，这是少有的「10B 也能本地跑」的开放权重底座——对机器人\u002F无人机开发者来说，终于不用再为「能塞进端侧」而妥协性能了。","https:\u002F\u002Fom-ai-lab.github.io\u002F2026_07_06_vlx_seek_1_5_zh.html","28b584de-85a2-4ef6-b4b5-511e1d9d5d73",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b543dbff-45f3-4888-839f-f6c39531b0b8","en","VLX-Seek 1.5: embodied perception in 0.6B\u002F3B\u002F10B, fewer illusions","Om AI Lianhui released VLX-Seek 1.5 on July 6 — a fine-grained perception VLM new version for on-device embodied scenarios. The keyword for this update is not \"bigger\" but \"more deployable\". The new version plans 0.6B \u002F 3B \u002F 10B three-tier model series, so that drones, robot dogs, surveillance cameras, and other terminals with different compute budgets can pick a suitable version — a rare engineering awareness in the era of mindlessly scaling VLMs that ultimately can't fit on-device. The architecture introduces more Linear Attention layers and faster OPN candidate region generation, with inference latency more friendly to the device. More noteworthy is the handling of \"object hallucination\". In embodied scenarios, the cost of a robot incorrectly tracking a non-existent target is far greater than missing one. VLX-Seek 1.5 introduces an explicit None output format: when the user asks about A and B in the image, if B doesn't exist, the model must output A's coordinates + B's None. On three benchmarks (HumanRef, VisDrone, RefDrone), the Object Hallucination metric (FP \u002F GT count) is lower than both the previous version and LocateAnything. On visual capability, the new training data adds more drone, surveillance, and robot viewpoints, and the auxiliary vision tower is also upgraded. On COCO, LVIS, RefCOCO, VisDrone, RefDrone, EmbSpatialBench and other benchmarks, VLX-Seek 1.5-3B surpasses several larger open-source\u002Fclosed-source VLMs. Om AI Lianhui announced the open-sourcing of the 10B version. In the embodied-perception track, this is a rare \"10B can also run locally\" open-weight foundation — for robot\u002Fdrone developers, no longer having to compromise performance just to \"fit on device\" is finally a reality.","vlx-seek-1-5","2026-07-06T02:00:00Z","2026-07-13T04:09:17.832856Z","2026-08-19T02:08:40.142862Z",true,"agent",308,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b6b9f5c8-0d71-4288-8782-0284fccfca8f","商汤 SenseNova-Vision：把「检测\u002F分割\u002F深度估计」统统塞进同一个生成式多模态基座","sensetime-sensenova-vision","2026-07-08T10:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f9bf6e21-2e8a-4571-ab7d-a4dba727b72a","ViiTorVoice-NAR：把 TTS 的「改一句重录」变成「改一词局部合成」","viitor-voice-nar-local-tts","2026-07-02T14:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]