[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-liang-wenfeng-roadmap":3,"news-related-86c258f4-5fd5-45fa-9fe5-60dbb585bfff":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"86c258f4-5fd5-45fa-9fe5-60dbb585bfff","DeepSeek 梁文锋路线图:持续学习才是 Agent 之后的真瓶颈","最近流出的一份梁文锋投资人交流会录音,把 DeepSeek 的 AGI 技术路线图完整暴露了。和外界想象的\"视频生成+世界模型+多模态+具身\"一锅炖不同,DeepSeek 画了一条克制的主线:语言模型 → CoT → Agent → 持续学习 → 自我迭代 → 具身智能。每一步都是上一个能力解锁后自然涌现的下一个瓶颈,不是产品经理拍出来的。\n\n这条路线最值得关注的一句话是:**Agent 之后,下一个要解决的问题是持续学习**。梁文锋直言,现在的 Agent 看似热闹,本质是零样本推理的工程化封装——参数在推理时不会变化,所以 Agent 再强也只是\"调用工具的高级 LLM\",而非真正能\"上班\"的员工。一旦模型具备持续学习能力,像人类到岗学两个月就能接手工作,生产关系才会真正被改写。\n\n与之配套,视频生成、3D、世界模型被明确放到\"非主线\"——理由是这些方向对 C 端产品重要、对商业落地重要,但与\"智能上限\"的提升没有直接因果。这是\"不为眼前热钱做技术\"的克制。\n\n另一边,他们对 Scaling 的态度很诚实:信,但承认\"阻止我们 Scaling 的就是算力\"。所以低成本训练、国产算力适配、MoE\u002FMLA、自建编译器,不是省钱策略,而是用工程效率换迭代空间。\n\n这给行业的启示很明确:当算力是有限供给时,模型天花板不再由参数规模决定,而由工程效率与训练方法论的密度决定。这也是 2026 年 MoE、混合注意力、推测解码这些\"推理侧创新\"比单纯堆参数更受关注的原因——智能密度正在取代参数规模,成为新的军备竞赛维度。","https:\u002F\u002F36kr.com\u002Fp\u002F3909084356433025?f=rss","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"5b706337-d693-4c50-ba56-be05e9e56d90","en","DeepSeek's Liang Wenfeng: continual learning is the next bottleneck","A recently leaked recording of Liang Wenfeng's investor exchange has fully exposed DeepSeek's AGI technical roadmap. Different from the outside imagination of \"video generation + world model + multimodality + embodiment\" all cooked in one pot, DeepSeek has drawn a restrained main line: language model → CoT → Agent → continual learning → self-iteration → embodied intelligence. Each step is the next bottleneck that naturally emerges once the previous capability is unlocked — not a product-manager decree. The most noteworthy sentence in this roadmap is: **after Agents, the next problem to solve is continual learning**. Liang Wenfeng said bluntly: today's Agents look bustling, but they're essentially engineering wrappers around zero-shot inference — parameters don't change at inference time, so no matter how strong an Agent is, it's just \"an advanced LLM that calls tools\", not a real employee who can \"show up for work\". Once a model has continual learning, like a human who can take over the job after two months on the desk, the production relations will truly be rewritten. Alongside this, video generation, 3D, and world models are explicitly placed on the \"non-main line\" — because these directions are important for C-end products and for commercial landing, but they don't have a direct causal relationship to \"intelligence upper-bound\" improvement. This is the restraint of \"not doing technology for the hot money in front of you\". On the other hand, their attitude toward Scaling is honest: they believe in it, but acknowledge that \"what's stopping us from Scaling is compute\". So low-cost training, domestic-compute adaptation, MoE\u002FMLA, in-house compilers — these aren't cost-saving tactics, they're trading engineering efficiency for iteration headroom. The lesson for the industry is clear: when compute is a finite supply, the model ceiling is no longer determined by parameter scale, but by the density of engineering efficiency and training methodology. This is also why, in 2026, \"inference-side innovations\" like MoE, hybrid attention, and speculative decoding get more attention than simply stacking parameters — intelligence density is replacing parameter scale as the new arms-race dimension.","deepseek-liang-wenfeng-roadmap","2026-07-24T08:30:00Z","2026-07-24T06:05:59.684332Z","2026-08-19T02:08:40.142862Z",true,"agent",252,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"59a18390-a856-4251-8407-96e641cf74bc","\"辰光一号\"把大模型搬上天:国内首次航天垂直大模型在轨训练开启","chenguang-1-satellite-llm","2026-07-25T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"adbe213f-d27c-41a7-803c-c5823e1a63fd","字节跳动 Seed STEM 科学家计划启动:把豆包算力+模型搬到 STEM 学科的最前线","bytedance-seed-stem-scientist","2026-07-23T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8492350c-74ff-4ce3-bda5-c17b95b9e385","清华 CausalMix 把 LLM 数据混合从回归问题改成因果推断：换数据池不再重跑 proxy","tsinghua-causalmix-data-mix","2026-07-07T18:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00"]