Tencent Hunyuan's multimodal understanding lead Hu Han recently left to start his own venture; his successor Tian Yonglong (former OpenAI researcher, MIT PhD) will take charge of the vision-language model direction. The original research group has been redirected to "world model" frontier research — another resource reshuffle after Yao Shunyu took over the large language model department. On the surface it looks like a personnel change, but at its core it's Tencent's judgment that the technology dividend from multimodal understanding has topped out. Hunyuan's internal researchers reveal: in image recognition, image captioning scenarios, the recognition accuracy of text, image, and video has already exceeded 85%, and the marginal return from piling on more data is rapidly diminishing; the harder visual reasoning depends on the language model's reasoning capability, and that path doesn't directly connect with multimodal research. On the commercial side, image-recognition tools can't find a paid scenario, while the truly high-value user demands — PPT, research reports, document processing — correspond to Agentic and Coding capabilities. The compute ledger is also pressuring the decision. Tencent's FY2025 capex was ¥79.2 billion, while Alibaba is ¥126 billion for the year and ByteDance is planning up to $70 billion — Tencent is at the bottom of the three. The research group must concentrate resources on the more "future" direction, and the world model happens to be at the intersection of three hot zones: NVIDIA Cosmos, Alibaba HappyOyster, and ByteDance Seed world model. Yao Shunyu's strategy is clear: base-model capability decides whether you can sit at the table, and the world model decides whether you can join the next round. Stepping back from multimodal looks like a contraction, but it's actually reloading the chamber.