[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hunyuan-exits-multimodal":3,"news-related-8452628f-d58e-417d-a340-6cafe1b75473":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8452628f-d58e-417d-a340-6cafe1b75473","腾讯混元撤出多模态理解,把子弹押给世界模型","腾讯混元多模态理解负责人胡瀚近日离职创业,接任者田永龙(前 OpenAI 研究员、MIT 博士)将负责视觉语言模型方向。原研究组被重新指向「世界模型」前沿研究,这是姚顺雨接手大语言模型部后的又一次资源重组。\n\n表面看是一次人事变动,本质是腾讯对多模态理解技术红利见顶的判断。混元内部研究员透露:识图、看图写话场景下,文字、图像、视频的识别准确率已经达到 85% 以上,继续堆数据带来的边际收益快速递减;更难的视觉推理要靠语言模型的推理能力,这条路又和多模态研究并不直接打通。商业端,识图类工具找不到付费场景,而 PPT、研报、文档处理这类真正高价值的用户需求,对应的是 Agentic 与 Coding 能力。\n\n算力账本同时在压决策。腾讯 2025 财年资本开支 792 亿元,而阿里单年 1260 亿元、字节计划最高 700 亿美元——三家里腾讯垫底。研究组必须把资源集中在更「未来」的方向,世界模型恰好同时撞上 NVIDIA Cosmos、阿里 HappyOyster、字节 Seed 世界模型的三重热门区。\n\n姚顺雨的策略清晰:基模能力决定能不能上牌桌,而世界模型决定下一轮能不能入局。从多模态退一步,看似收缩,其实是把子弹重新装膛。","https:\u002F\u002F36kr.com\u002Fp\u002F3907934819521670","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"95d7995a-fddb-47ba-b8e6-e976ac65414b","strategy",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"f7515eef-b40d-459b-bbdc-3c987f618119","en","Tencent Hunyuan exits multimodal, bets on world models","Tencent Hunyuan's multimodal understanding lead Hu Han recently left to start his own venture; his successor Tian Yonglong (former OpenAI researcher, MIT PhD) will take charge of the vision-language model direction. The original research group has been redirected to \"world model\" frontier research — another resource reshuffle after Yao Shunyu took over the large language model department. On the surface it looks like a personnel change, but at its core it's Tencent's judgment that the technology dividend from multimodal understanding has topped out. Hunyuan's internal researchers reveal: in image recognition, image captioning scenarios, the recognition accuracy of text, image, and video has already exceeded 85%, and the marginal return from piling on more data is rapidly diminishing; the harder visual reasoning depends on the language model's reasoning capability, and that path doesn't directly connect with multimodal research. On the commercial side, image-recognition tools can't find a paid scenario, while the truly high-value user demands — PPT, research reports, document processing — correspond to Agentic and Coding capabilities. The compute ledger is also pressuring the decision. Tencent's FY2025 capex was ¥79.2 billion, while Alibaba is ¥126 billion for the year and ByteDance is planning up to $70 billion — Tencent is at the bottom of the three. The research group must concentrate resources on the more \"future\" direction, and the world model happens to be at the intersection of three hot zones: NVIDIA Cosmos, Alibaba HappyOyster, and ByteDance Seed world model. Yao Shunyu's strategy is clear: base-model capability decides whether you can sit at the table, and the world model decides whether you can join the next round. Stepping back from multimodal looks like a contraction, but it's actually reloading the chamber.","tencent-hunyuan-exits-multimodal","2026-07-23T00:07:00Z","2026-07-23T18:02:50.690926Z","2026-08-19T02:08:40.142862Z",true,"agent",87,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"faab4a6c-9cb0-4f2a-a5bf-1f122306008b","Wan-Streamer v0.2：分辨率 192p→640p，保住 200ms","alibaba-wan-streamer-v0-2","2026-07-10T16:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4c951e48-08c9-4d5d-b9a0-3dfdd1b04bed","Visics 把 Object Trajectory 做成统一中间表征：通用具身大模型有了自己的 Token","visics-vloa-object-trajectory-embodied","2026-06-25T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"0896ec0a-c4af-4e36-a357-de4610c57e97","微信AI助手「小微」灰度内测：WeLM + DeepSeek 双模型架构，14.32亿月活的 Agent 落地实验","wechat-xiaowei-welm-deepseek-agent","2026-06-24T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"efa5f558-e27f-4889-bf71-74cbef874ace","即梦 AI 把 Seedance 2.0 推上原生 4K:把超分这道工序从后期流水线里拿掉","seedance-2-0-native-4k-jimeng","2026-06-24T08:00:00+00:00"]