[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sensetime-sensenova-vision":3,"news-related-b6b9f5c8-0d71-4288-8782-0284fccfca8f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b6b9f5c8-0d71-4288-8782-0284fccfca8f","商汤 SenseNova-Vision：把「检测\u002F分割\u002F深度估计」统统塞进同一个生成式多模态基座","商汤科技联合上海人工智能实验室(S-Lab)、南洋理工大学和香港中文大学,在 arXiv 上推出 SenseNova-Vision,把计算机视觉任务重新定义为「统一多模态生成」问题。\n\n传统 CV 模型要为检测、分割、深度估计、关键点等不同任务分别设计专用预测头,而 SenseNova-Vision 用一个统一多模态基座,通过自然语言指令加可选视觉提示,直接生成文本(符号输出)、图像(密集空间预测)或图文混合(组合任务)。论文显示,这一单一模型在结构化视觉理解、密集几何预测、分割、多视图几何等任务上均可对标专用系统。\n\n更关键的是,团队同步开源了 SenseNova-Vision Corpus —— 一个跨文本、图像和混合目标的视觉指令-响应语料库,以及配套的预训练权重(GitHub: OpenSenseNova\u002FSenseNova-Vision)。CV 社区第一次可以像用 LLM 一样,用一个基座替换整套视觉工具箱。\n\n这条路线对工程界的意义远大于一个 SOTA 数字:它把「视觉能力即文本生成」的范式推到工程级,未来通用基座不必再外挂 YOLO、Segment Anything、Depth Anything 各自的小模型 —— 一个生成式基座端到端处理。这可能是 CV 行业从「任务驱动」转向「生成驱动」的下一道分水岭。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.06560","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"29840a27-c35e-4e5f-b4fc-60730e5f0ebd","en","SenseNova-Vision: detection, segmentation, depth in one base","SenseTime, together with Shanghai AI Lab (S-Lab), Nanyang Technological University, and the Chinese University of Hong Kong, launches SenseNova-Vision on arXiv, redefining computer vision tasks as a \"unified multimodal generation\" problem. Traditional CV models need separate dedicated prediction heads for different tasks like detection, segmentation, depth estimation, and keypoints, while SenseNova-Vision uses a unified multimodal base, with natural language instructions plus optional visual prompts, directly generating text (symbol output), images (dense spatial prediction), or text-image hybrid (composite tasks). The paper shows that this single model can be benchmarked against dedicated systems on structured visual understanding, dense geometric prediction, segmentation, multi-view geometry, and other tasks. More critically, the team open-sourced SenseNova-Vision Corpus simultaneously — a vision instruction-response corpus spanning text, image, and mixed targets, along with the matching pretrained weights (GitHub: OpenSenseNova\u002FSenseNova-Vision). The CV community can, for the first time, use a base to replace the entire vision toolbox just like using an LLM. The significance of this path for the engineering world is far greater than a single SOTA number: it pushes the \"vision capability is text generation\" paradigm to engineering-grade, future general bases no longer need to plug in dedicated small models like YOLO, Segment Anything, Depth Anything — a generative base handles it end-to-end. This may be the next watershed for the CV industry shifting from \"task-driven\" to \"generation-driven\".","sensetime-sensenova-vision","2026-07-08T10:15:00Z","2026-07-08T10:06:54.338105Z","2026-08-19T02:08:40.142862Z",true,"agent",147,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a36d9d97-42de-4c87-88e9-cdc173b9ab4b","VLX-Seek 1.5 把端侧具身感知切成 0.6B\u002F3B\u002F10B 三档：用 None 输出压住目标幻觉","vlx-seek-1-5","2026-07-06T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f9bf6e21-2e8a-4571-ab7d-a4dba727b72a","ViiTorVoice-NAR：把 TTS 的「改一句重录」变成「改一词局部合成」","viitor-voice-nar-local-tts","2026-07-02T14:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]