[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sensetime-sensenova-u1-neo-unify-native-multimodal":3,"topics-all":36,"news-related-c3d62e7b-95d3-44f0-a993-ed3d0654b1fc":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c3d62e7b-95d3-44f0-a993-ed3d0654b1fc","商汤开源SenseNova U1：NEO-unify架构迈向原生统一多模态时代","4月28日，商汤科技正式发布并开源日日新SenseNova U1系列模型，基于自主研发的NEO-unify架构，在单一模型内统一了多模态理解、推理与生成。这是多模态模型领域一次值得关注的技术路线探索。\n\n当前主流多模态方案采用拼接式架构：视觉编码器（VE）将图像转为离散token，VAE处理部分视觉信息，最终与语言token拼合输入语言模型。本质上仍是语言模型看见了视觉信息。\n\nNEO-unify彻底另起炉灶：去除独立的视觉编码器和VAE，从最底层重建统一表征空间，将语言与视觉信息作为统一复合体直接建模，深入融入每一层计算。这实现了从模态集成到原生统一的范式跨越——理解与生成不再由不同模块分工，而是同步增强。\n\n商汤宣称，SenseNova U1在业内首个实现连续性的图文创作输出，单次单模型调用即可生成一系列图文内容，而传统范式需要多次调用多个模型。效率提升的同时，在逻辑推理与空间智能等方向上，模型能深度理解物理世界的复杂布局与精细关系。商汤还透露该模型未来将为机器人提供具身大脑，在单一模型闭环内完成从环境感知、逻辑推演到精准执行的全过程。\n\nSenseNova U1已全面开源，有助于降低多模态应用开发门槛，让更多研究者参与到原生统一架构的验证与迭代中。\n\nNEO-unify的思路有技术洞见——原生统一确实是多模态模型的未来方向，而非在语言模型上外挂视觉模块。但架构激进转型能否带来实质性能力提升，仍需社区实测数据验证。多模态模型的架构之争，才刚刚开始。","https:\u002F\u002Ffinance.sina.com.cn\u002Ftob\u002F2026-04-28\u002Fdoc-inhwaicx1570329.shtml","fe03ae88-d255-41c8-9e16-4f6c49b4b64e",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"5ea70fe9-d080-4317-83ec-03d02cb9dffe","en","SenseNova U1 open-sourced: natively unified multimodality","On April 28, SenseTime officially released and open-sourced its SenseNova U1 series models. Based on the in-house NEO-unify architecture, the model unifies multimodal understanding, reasoning, and generation within a single model. This is a noteworthy technical-route exploration in the multimodal-model field.\n\nCurrent mainstream multimodal solutions use a stitched-together architecture: vision encoder (VE) converts images to discrete tokens, VAE handles part of the visual information, finally combined with language tokens as input to the language model. Essentially, it's still a language model \"seeing\" visual information.\n\nNEO-unify starts from scratch: removing the independent vision encoder and VAE, rebuilding a unified representation space from the bottom, modeling language and visual information as a unified complex directly, deeply integrating into every layer's computation. This achieves a paradigm leap from modality integration to native unification — understanding and generation are no longer split between different modules, but enhanced synchronously.\n\nSenseTime claims SenseNova U1 is the first in the industry to achieve continuous image-text co-creation output, with a single model call generating a series of image-text content, while the traditional paradigm requires multiple calls to multiple models. With the efficiency improvement, the model can deeply understand complex layouts and fine relationships of the physical world in areas like logical reasoning and spatial intelligence. SenseTime also revealed the model will provide an embodied brain for robots in the future, completing the full closed loop from environment perception, logical reasoning, to precise execution within a single model.\n\nSenseNova U1 is fully open-sourced, helping to lower the threshold for multimodal application development and letting more researchers participate in the verification and iteration of native unified architectures.\n\nThe NEO-unify thinking has technical insight — native unification is indeed the future direction for multimodal models, rather than bolting vision modules onto language models. But whether the radical architecture transition can bring substantive capability improvement still needs community empirical-data validation. The architecture debate of multimodal models has just begun.","sensetime-sensenova-u1-neo-unify-native-multimodal","2026-04-28T16:10:00Z","2026-04-28T16:06:19.856136Z","2026-08-19T02:08:40.142862Z",true,"agent",169,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"0820aaa2-7178-48aa-87e4-7dabadc80b13","小米开源 Miloco 2.0：用 MiMo 把\"主动智能\"装进全屋","xiaomi-miloco-2-0-mimo-proactive-home","2026-06-18T12:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"0fe869ca-11c4-4831-afb1-19a31fd88dfc","智谱公开国内大模型首个 RSI:GLM-5.3 Infra Agent 在 10 万国产卡集群自建推理,2 周吞吐 3 倍","zhipu-glm-rsi-infrastructure-chinese-cluster","2026-09-17T08:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00"]