[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-baai-wang-zhongyuan-2012-moment-latent":3,"news-related-8e55d420-7ed9-4269-b18c-6000e765ea5e":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8e55d420-7ed9-4269-b18c-6000e765ea5e","智源王仲远：世界模型处在\"2012 时刻\"，\"潜空间统一\"是中美同跑的第五种解法","智源研究院（BAAI）院长王仲远在硬氪专访中，把\"世界模型\"（World Model）拆成四条分岔路，并抛出\"中国与海外正处于同一起跑线\"的判断。\n\n他归纳的四类主流路线分别是：以语言为中心（VLM、VLA，文本空间预测下一个词但学不到物理后果）、以像素为中心（Sora、Seedance 等视频生成类，像素级画面但不懂因果）、以三维结构为中心（World Labs Marble 等，几何结构≠物理状态），以及以视觉表征为中心（LeCun 的 JEPA 系列，预测表征压缩而非物理规律）。\n\n智源走的是\"第五种\"路：把所有模态压缩进同一潜空间（Latent Space），再由不同 Decoder 按需还原成画面、动作、位置。这套\"潜空间统一表征\"已陆续接入悟界·Emu3\u002FEmu3.5、悟·Physis 和悟界·RoboBrain Orca；其中 Emu3.5 验证了\"类 LLM 的统一架构在多模态上能 Scale Up\"，给世界模型阶段的基础设施铺路。\n\n王仲远明确表态\"视频生成不等于世界模型\"。OpenAI 当年用 World Simulator 描述 Sora 让这个词被泛化，但真正的世界模型核心是\"下一个物理状态预测\"（NSP），物理正确、动作因果可溯、长时序一致、跨场景泛化是四项硬指标。\n\n在他看来，世界模型还处在\"深度学习的 2012 年前后\"，数据孤岛、路线未定、Benchmark 还在打架，距离 ChatGPT 时刻还有 3 到 5 年。但和 LLM 时代不同，\"中美没有差距\"。短期 VLA 仍是工厂分拣等场景主力；长期看，能预测物理状态、指挥机器人决策的世界基座模型，才是 AGI 进入物理世界的真正底座。","https:\u002F\u002F36kr.com\u002Fp\u002F3853016586359817","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e4dfcd8d-7926-40cd-a1f2-4e7439b1f109","en","BAAI's Wang Zhongyuan: world models at their 2012 moment","36Kr's feature on BAAI (Beijing Academy of Artificial Intelligence) researcher Wang Zhongyuan argues that world models are at a \"2012 moment\" — comparable to deep learning's breakthrough around 2012. The argument: the \"latent space unification\" approach is the fifth major paradigm in AI, and the US and China are racing toward it together.\n\nThe \"2012 moment\" framing: Wang draws a parallel to 2012, when AlexNet won ImageNet and triggered the deep learning revolution. World models in 2026 are at a similar inflection point — the recent breakthroughs (Genie 3, Sora 2, Wan-Streamer) suggest that \"world model\" is the next general-purpose AI paradigm.\n\nThe \"latent space unification\" insight: the \"fifth answer\" is the idea that all modalities (text, image, video, audio, action) can be unified in a single latent space, and a single model can operate over this unified space. This is a more ambitious version of the \"multimodal\" idea — instead of having separate encoders for each modality, a single latent space captures everything.\n\nThe \"US-China race\" angle: Wang argues that the US and China are both pursuing latent space unification, but with different approaches. The US labs (OpenAI, Anthropic, Google) are pursuing it via scaling (bigger models, more data). The Chinese labs (BAAI, Qwen, DeepSeek) are pursuing it via architecture (MoE, SSM, novel attention). The two approaches will likely converge.\n\nThe bigger takeaway: \"latent space unification\" is the next major AI paradigm. The \"2012 moment\" framing suggests we're on the cusp of a major transformation, and the \"latent space unification\" approach is the most promising direction. For the industry, this signals that \"AI as a unified latent space\" is the long-term vision, and the next 5-10 years will be defined by progress toward this goal.","baai-wang-zhongyuan-2012-moment-latent","2026-06-15T10:00:00Z","2026-06-15T10:14:24.251929Z","2026-08-19T02:08:40.142862Z",true,"agent",125,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"0896ec0a-c4af-4e36-a357-de4610c57e97","微信AI助手「小微」灰度内测：WeLM + DeepSeek 双模型架构，14.32亿月活的 Agent 落地实验","wechat-xiaowei-welm-deepseek-agent","2026-06-24T10:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"0f0af70c-3e53-4d6a-b49c-9cd0aef71e73","映界科技把 4D 高斯和 VLM 拼成空间记忆：给具身机器人补一块可被 LLM 查询的感知层","mirrorspace-yingjie-4d-gaussian-vlm-embodied","2026-06-22T22:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]