[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-shai-lab-intern-s2-preview-397b":3,"news-related-a16a4f36-cb00-4f14-a394-80d49075a323":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a16a4f36-cb00-4f14-a394-80d49075a323","上海AI Lab 发布 Intern-S2-Preview-397B：把「记忆」与「思考」拆开，397B 跑出万亿模型效果","WAIC 2026 开幕当天，上海人工智能实验室发布书生系列新成员 Intern-S2-Preview-397B。这不是又一份参数堆料，而是一次底座架构级转向——放弃「一切都在 Transformer 里」的传统路径，把「知识承载」与「推理计算」拆成两条独立但可协同的引擎。\n\n新架构核心是一对组件：**Memory Decoder** 把专业知识做成可插拔外部记忆模块，按需挂载到基座；**Mobius** 是全新推理主干，通过反向残差连接让深层隐状态访问浅层知识，用动态隐空间推理替代 Token。结果相当硬：在分子设计、材料结构生成等科学任务上，397B 的 Intern-S2-Preview 追平了实验室此前的万亿参数模型，端到端推理效率提升约 4 倍。\n\n配套 **InternBootcamp** 把电路设计、金融建模等真实任务变成「行动—反馈」式交互场景，让模型在试错中内化领域逻辑；「书生·端砚」已落地生命科学、关键材料、半导体、核聚变、量子、地球气象六大领域。当参数扩张撞上算力与能耗天花板，「记忆—推理解耦 + 任务级 RL」给了科学智能体一条不靠纯堆参数也能往前走的样本。","https:\u002F\u002Fwww.qbitai.com\u002F2026\u002F07\u002F452942.html","3bd971a8-3897-43d9-84ac-43879efd2f94",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"40a3c87a-8253-4059-a0df-31c7be6858b4","en","Intern-S2-397B splits memory from thinking, hits trillion-scale","On the opening day of WAIC 2026, Shanghai AI Lab released a new member of the Shusheng series, Intern-S2-Preview-397B. This is not another parameter-stacking exercise, but an architecture-level turn — abandoning the \"everything lives inside the Transformer\" path, and splitting \"knowledge bearing\" from \"reasoning computation\" into two independent but collaborative engines. The new architecture's core is a pair of components: **Memory Decoder** turns domain knowledge into pluggable external memory modules, mounted on demand onto the base; **Mobius** is a brand-new reasoning backbone that uses reverse residual connections so deep hidden states can access shallow-layer knowledge, replacing tokens with dynamic latent-space reasoning. The results are tough: on science tasks like molecular design and material structure generation, Intern-S2-Preview-397B matches the lab's previous trillion-parameter model, with end-to-end inference efficiency up roughly 4×. The companion **InternBootcamp** turns circuit design, financial modeling and other real tasks into \"action-feedback\" interactive scenarios, letting the model internalize domain logic through trial and error; \"Shusheng · Duanyan\" has already landed in six fields: life science, key materials, semiconductors, nuclear fusion, quantum, and Earth weather. When parameter expansion hits compute and energy ceilings, \"memory-reasoning decoupling + task-level RL\" gives scientific agents a sample path that doesn't rely purely on stacking parameters to keep moving forward.","shai-lab-intern-s2-preview-397b","2026-07-18T02:00:00Z","2026-07-18T02:16:57.341954Z","2026-08-19T02:08:40.142862Z",true,"agent",273,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d1e8997e-bb60-453d-9ef8-71b8bdde5386","Harvey 首个自研法律模型 Tenet 曝光:底座没选 GPT 和 Claude,选了 Kimi K3","harvey-tenet-kimi-k3-legal-model","2026-08-18T17:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00"]