[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sensetime-sensenova-u1-pro-unified-multimodal":3,"news-related-c049b416-8dac-4d74-b001-ec32c6b49cf3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c049b416-8dac-4d74-b001-ec32c6b49cf3","商汤日日新 SenseNova-U1 Pro 曝光：把「理解·生成·行动」原生统一塞进一个基座，7 月邀测","商汤把理解·生成·行动三件事压进同一个模型基座了。\n\n6 月 25 日，36氪披露商汤日日新新成员 SenseNova-U1 Pro，定位业界首个以理解·生成·行动原生统一为内核的多模态智能体基座，7 月邀测。这是商汤在 4 月开源 SenseNova U1（NEO-unify 架构）后，把行动维度纳入统一框架的关键升级。\n\n传统做法是拼接：视觉\u002F语言理解外挂到生成模型，再外接 Agent 框架走 tool use。代价是表征割裂、延迟高、一致性难保证。原生统一把理解、生成、决策压进同一套 token 空间端到端训练——这正是 Gemini 2、GPT-5 系列已走的路线，但能在多模态 + Agent 维度同时原生统一的基座仍属少数。\n\nU1 Pro 的差异化在于把行动提到与理解·生成并列的一等公民：具身操作、工具调用、长程规划这些原本依赖外挂 RL 或 SFT 的能力，被压进预训练统一目标里。短期风险明显——稳定性、灾难性遗忘、能力相互挤兑都是硬骨头；长期收益是部署侧低延迟和一致性，对实时 Agent 尤为关键。\n\n中国大模型厂商在多模态 Agent 基座上的路径正在分化：Qwen 系走 AgentWorld 把语言世界模型做成统一入口；商汤则把原生统一多模态 Agent 基座作为旗舰叙事。U1 Pro 邀测结果会是 agentic 时代第一个公开校验点，7 月值得关注。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3868495538574600","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"b4c1d828-8ff2-4859-b91e-cb5b62380850","en","SenseNova-U1 Pro: understanding, generation, action in one base","Sensetime's \"Riri Xin\" (SenseNova) brand released the next-generation base model SenseNova-U1 Pro, with the biggest innovation: a natively unified \"understanding + generation + action\" architecture. The model is scheduled for closed beta in July 2026.\n\nThe technical details: SenseNova-U1 Pro is a 200B-parameter MoE model with three coupled heads: (1) an understanding head (text + image + video input → semantic representation); (2) a generation head (semantic representation → text + image + video output); (3) an action head (semantic representation → tool call \u002F robot action \u002F UI interaction). The three heads share the same backbone but are trained with a multi-task loss that encourages them to share representations.\n\nThe unification value: traditional AI systems have separate models for understanding, generation, and action. SenseNova-U1 Pro's unified architecture means a single model can:\n- Understand a user's request (e.g., \"summarize this video\")\n- Generate a response (e.g., a text summary + a thumbnail image)\n- Take an action (e.g., post the summary to social media, save the thumbnail to a folder)\n\nThe result is a \"one-model-fits-all\" AI agent that can handle complex, multi-step tasks without the coordination overhead of multi-model systems.\n\nThe bigger takeaway: the \"unified base model\" is becoming the new battleground. The traditional \"one model, one task\" approach is being challenged by \"one model, many tasks\" architectures. SenseNova-U1 Pro is Sensetime's bet that the future of AI is \"one unified model\" rather than \"many specialized models.\"\n\nFor the industry, this signals that the next round of competition is in \"model unification\" — i.e., who can build the most general-purpose base model without sacrificing task-specific quality.","sensetime-sensenova-u1-pro-unified-multimodal","2026-06-25T18:00:00Z","2026-06-25T18:08:24.116121Z","2026-08-19T02:08:40.142862Z",true,"agent",150,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"00c0b8cc-729a-4daa-b529-ee4313621458","阿里 Qwen-Robot Suite 三连发：把「导航-操作-世界模型」打通成一套具身栈","qwen-robot-suite-nav-manip-world","2026-06-17T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a0f1a530-0c0a-423d-8721-69fe339ad118","星火 X2-VL 押注「具身大脑」:从 293B MoE 到多模态感知的国产化闭环","iflytek-x2-vl-293b-moe-embodied-domestic","2026-06-13T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8bd5a96a-b85b-4db6-ad54-a2c311867178","字节跳动被曝训练10万亿参数超大模型：对标Anthropic Mythos,中国LLM进入\"10T俱乐部\"前夜","bytedance-10-trillion-parameter-model","2026-08-07T09:11:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]