[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-10t-mythos-zhangyiming-no-distill-2026-08":3},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":30,"news_slug":37,"published_at":38,"created_at":39,"modified_at":40,"is_published":41,"publish_type":42,"image_url":14,"view_count":43},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","据《金融时报》报道，字节跳动正在训练一款参数量约 10 万亿的大模型，仍处预训练阶段；该数字约为月之暗面 Kimi K3 的三倍多，按行业估算逼近 Anthropic Mythos 5（约 8 万亿）的体量。一天前，张一鸣罕见地内部表态拒绝用蒸馏换能力，为这场规模赛跑定下长期主义基调。","# 字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态\n\n8 月 7 日，《金融时报》报道字节跳动正在训练一款参数量约 10 万亿的大模型，目前仍处于预训练阶段。预训练通常需要 3 到 6 个月，之后才会进入微调与正式发布——也就是说这颗\"超大模型\"距离面世还有相当一段路要走（[联合早报\u002F金融时报转载](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328)）。\n\n报道披露了几个关键数字：凭借 10 万亿参数量，字节这款模型体量将是中国 AI 新锐月之暗面旗下 Kimi K3（拥有 2.8 万亿参数）的三倍有余；在此之前，美团的 LongCat-2.0 与 DeepSeek 的 V4-Pro 曾以 1.6 万亿的参数量领跑中国本土 AI 行业；阿里最新发布的 Qwen3.8-Max 约 2.4 万亿参数（[21财经\u002F金融界](https:\u002F\u002Fm.21jingji.com\u002Farticle\u002F20260807\u002Fherald\u002Febb1e261c9aba6a5959ca58fd3cd0b3d.html)）。按行业估算，Anthropic 最先进的 Mythos 5 参数量约为 8 万亿，Fable 5 约为 5 万亿，**这些数字均为行业估算，并非 Anthropic 公开披露**（[FourWeekMBA](https:\u002F\u002Ffourweekmba.com\u002Fai-bytedance-10-trillion-parameter-model-compute-sovereignty\u002F)）。\n\n## 「10 万亿」意味着什么：MoE 时代的参数游戏\n\n需要先拆解一个误解：报道中所有这些数字都是**总参数**，而非\"激活参数\"。在现代前沿模型稀疏 MoE（混合专家）架构下，10 万亿总参数意味着绝大多数权重在任意一次 token 前向上是\"睡着\"的。总参数更像一张声明——\"我买得起这份算力、堆得起这份数据\"，而不是\"我用得上这份能力\"。\n\n把这个视角放回国内市场：DeepSeek V4 系列、阿里 Qwen3.8-Max、月之暗面 Kimi K3 走的都是稀疏 MoE + 极致训练效率路线——用尽量小的激活参数量去逼近前沿；字节跳动这次则选择了另一端：在规模上直接堆到 10 万亿。两种打法的成本结构截然不同，前者押注\"单位推理成本下的能力密度\"，后者押注\"绝对能力上限\"。\n\n## 张一鸣前一天的内部表态：拒绝把蒸馏当捷径\n\n巧合的是，8 月 6 日《澎湃新闻》已经披露字节跳动创始人张一鸣在 7 月 Seed 团队全员大会上的罕见表态：\"字节跳动不会把蒸馏当作提升 AI 模型能力的捷径，即便这意味着公司眼下落后于国内竞争对手\"（[联合早报转载](https:\u002F\u002Fwww.zaobao.com.sg\u002Fnews\u002Fchina\u002Fstory20260806-9481691)）。\n\n张一鸣说\"做模型要坚持长期主义、延迟满足感，而不是用别人的输出换一时的榜单排名\"。报道还指出，字节跳动内部对开源模型严禁蒸馏，甚至通过 API 检测等方式在内部加强相关限制——因为张一鸣认为蒸馏会\"干扰真正意义上的长期技术突破\"。\n\n把这两条新闻摆在一起读，10 万亿参数的赌注就有了更清晰的注脚：**字节跳动选择不抄近道，于是必须砸远路**。在 DeepSeek 把开源能力天花板一次次抬高、月之暗面用 Kimi K3 拿下 57 分 Artificial Analysis Intelligence Index 的压力下，字节跳动没有走蒸馏这条性价比最高的能力迁移路径，转而选择\"用绝对规模换代差\"——这条路更贵、更慢，但更难被复制。\n\n## 路线之争：中国 AI 第一次同时跑两个方向\n\n更深一层看，10 万亿参数这件事真正值得关注的，不是\"中国能不能追上美国前沿\"，而是\"中国内部同时跑出了两条互斥路线\"。\n\n一条是 DeepSeek 主张的\"小而省\"——V4 Flash 用 284B 总参数\u002F13B 激活参数的 MoE，在 Artificial Analysis Intelligence Index 拿到 50 分，混合价格约 0.06 美元\u002F百万 token，把 GPT-5.6 Luna 拉下约 65% 单任务成本。另一条是字节跳动主张的\"大而猛\"——把总参数堆到 Mythos 5 同档（甚至按估算略高），用算力密度换能力上限。\n\n两条路线的成本曲线、目标客户、商业模型都不一样。在 DeepSeek 用开源权重持续给 API 价格施压的环境里，字节跳动押注的是另一类客户：愿意为\"绝对能力上限\"付溢价的 Agent、复杂推理、企业级 RAG 场景。这也是为什么张一鸣那句\"愿意为长期目标牺牲一部分短期收益\"必须配上 10 万亿参数才听得懂——没有这个规模的算力承诺，那句表态就只是漂亮话。\n\n## 算力主权视角：出口管制可能正在失效\n\n另一个无法回避的视角是算力。10 万亿参数的总规模在预训练阶段的算力消耗，已经逼近美国出口管制原本试图限制的能力区间。FourWeekMBA 注意到这一点：如果字节跳动真的在用 10 万亿总参数做 MoE 预训练，那意味着中国实验室已经在某种程度上绕开了对高端 NVIDIA 加速器的依赖——无论是控制收紧前的存量芯片、华为昇腾等国产替代、还是训练效率提升带来的对冲。\n\n但这里必须加几个保留：**字节跳动未公开证实 10 万亿这个数字**，《金融时报》的来源也只是\"知情人士\"；Mythos 5 与 Fable 5 的参数是行业估算，并非官方披露；模型最终是否会发布、是否仍是 MoE 架构、激活参数多少，都还是未知数。10 万亿这个数字今天的最大价值，是它对资本、监管、竞争对手发出的信号——而不是对用户发出的能力承诺。\n\n## 所以呢\n\n对中国 AI 行业来说，下半年要回答的问题已经从\"哪家模型最强\"变成了\"规模与效率哪条路线能跑出商业模型\"。DeepSeek 路线用开源权重在压价格下限，字节跳动路线用绝对规模在顶能力上限。10 万亿参数是否最终跑通还是其次——**它证明的是中国头部公司愿意为前沿算力持续买单**，这件事本身就把全球 AI 竞争的格局往前推了一步。对从业者来说，比关心\"什么时候能调到这颗模型\"更值得想清楚的是：当模型层的差异化不再来自参数大小而是来自训练方法、后训练配方、推理成本控制时，自己的产品壁垒到底在哪里。","https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260806-9481691","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21,24,27],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":28,"name":29,"slug":29,"description":14,"color":14},"95d7995a-fddb-47ba-b8e6-e976ac65414b","strategy",[31],{"id":32,"lang":33,"title":34,"summary":35,"content":36},"dba7d558-9df0-4c05-a3ff-c7c94b7ce42d","en","ByteDance's 10-Trillion-Parameter Bet: Scale Racing Meets Zhang Yiming's \"No-Distillation\" Stance","According to the Financial Times, ByteDance is training an AI model with around 10 trillion parameters, still in the pre-training stage. The figure is more than three times Moonshot's Kimi K3 (2.8T) and approaches the estimated ~8T parameter count of Anthropic's Mythos 5 by industry estimates. One day earlier, founder Zhang Yiming made the rare internal declaration that ByteDance would not use distillation as a shortcut for capability gains, anchoring the long-term-ist framing of this scale bet.","# ByteDance's 10-Trillion-Parameter Bet: Scale Racing Meets Zhang Yiming's \"No-Distillation\" Stance\n\nOn August 7, the Financial Times reported that ByteDance is training an AI model with around 10 trillion parameters, currently in the pre-training stage. Pre-training typically takes 3 to 6 months, followed by fine-tuning and release — meaning this \"extra-large model\" is still a long way from public availability ([Lianhe Zaobao \u002F FT reprint](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328)).\n\nThe report disclosed several key numbers: at 10T parameters, ByteDance's model is more than three times the size of Moonshot's Kimi K3 (2.8T). Previously, Meituan's LongCat-2.0 and DeepSeek's V4-Pro led the Chinese frontier at 1.6T; Alibaba's latest Qwen3.8-Max sits at around 2.4T ([21 Jingji \u002F JRJ](https:\u002F\u002Fm.21jingji.com\u002Farticle\u002F20260807\u002Fherald\u002Febb1e261c9aba6a5959ca58fd3cd0b3d.html)). By industry estimates, Anthropic's most advanced Mythos 5 carries around 8T parameters, Fable 5 around 5T — **these figures are industry estimates, not publicly confirmed by Anthropic** ([FourWeekMBA](https:\u002F\u002Ffourweekmba.com\u002Fai-bytedance-10-trillion-parameter-model-compute-sovereignty\u002F)).\n\n## What \"10 Trillion\" Actually Means: Parameter Games in the MoE Era\n\nA necessary clarification: all these numbers refer to **total parameters**, not \"active parameters.\" In modern frontier sparse MoE architectures, 10T total parameters means the vast majority of weights are \"asleep\" during any single token forward pass. Total parameter count is closer to a statement — \"I can afford this compute, I can stack this much data\" — rather than a capability claim.\n\nPut back in the domestic context: DeepSeek V4, Alibaba Qwen3.8-Max, and Moonshot Kimi K3 all pursue sparse MoE + extreme training efficiency — pushing frontier capability with as few active parameters as possible. ByteDance has now chosen the opposite end: stacking total parameters up to 10T. The cost structures diverge sharply — one bets on \"capability density per unit inference cost,\" the other on \"absolute capability ceiling.\"\n\n## Zhang Yiming's Day-Before Internal Statement: Refusing Distillation as a Shortcut\n\nCoincidentally, on August 6, The Paper reported Zhang Yiming's rare statement at ByteDance's July all-hands for the Seed team: \"ByteDance will not use distillation as a shortcut to lift AI model capability, even if that means we currently lag behind domestic competitors\" ([Lianhe Zaobao reprint](https:\u002F\u002Fwww.zaobao.com\u002Fsg\u002Fnews\u002Fchina\u002Fstory20260806-9481691)).\n\nZhang said \"making models demands long-termism and delayed gratification — not trading someone else's outputs for a momentary spot on a leaderboard.\" The report also noted ByteDance internally forbids distillation from open-source models, even enforcing limits via API detection internally — because Zhang believes distillation \"interferes with truly long-term technological breakthroughs.\"\n\nRead together, the 10T bet takes on a clearer annotation: **ByteDance chose not to take the shortcut, so it had to commit to the long road.** Under pressure from DeepSeek repeatedly raising the open-source capability ceiling and Moonshot's Kimi K3 hitting 57 on the Artificial Analysis Intelligence Index, ByteDance did not take distillation — the most cost-efficient path to capability migration — and instead chose \"use absolute scale to substitute for the gap.\" This road is more expensive, slower, but harder to copy.\n\n## The Route Debate: China AI Runs Both Directions at Once\n\nA deeper layer: the truly interesting thing about 10T parameters is not \"can China catch up with the US frontier,\" but \"China is internally running two mutually exclusive routes simultaneously.\"\n\nOne is DeepSeek's \"small and frugal\" thesis — V4 Flash uses 284B total \u002F 13B active MoE, scores 50 on Artificial Analysis Intelligence Index, with blended price around $0.06 per million tokens, ~65% lower single-task cost than GPT-5.6 Luna. The other is ByteDance's \"large and fierce\" bet — stacking total parameters into Mythos 5's tier (or slightly above by estimates), trading compute density for capability ceiling.\n\nThe two routes have different cost curves, target customers, and business models. In the environment where DeepSeek keeps applying price pressure via open-source weights, ByteDance is betting on a different customer class: those willing to pay a premium for \"absolute capability ceiling\" — Agent workloads, complex reasoning, enterprise-grade RAG. This is why Zhang Yiming's line \"willing to sacrifice some short-term returns for long-term goals\" only makes sense paired with a 10T compute commitment — without that scale commitment, the statement would just be a pretty phrase.\n\n## Compute Sovereignty Lens: Export Controls May Be Thinning\n\nA view that cannot be avoided is compute. The pre-training compute footprint of a 10T-parameter model already approaches the capability band US export controls were designed to constrain. FourWeekMBA flagged this point: if ByteDance is genuinely running MoE pre-training at 10T total parameters, it implies Chinese labs have found ways around dependence on cutting-edge NVIDIA accelerators — whether through stockpiled chips acquired before controls tightened, Huawei Ascend and other domestic substitutes, or training-efficiency improvements that offset the hardware gap.\n\nBut several caveats are required: **ByteDance has not publicly confirmed the 10T figure**, the FT source is \"people familiar with the plan\"; Mythos 5 and Fable 5 parameter counts are industry estimates, not official disclosures; whether the model will ship at all, whether it remains MoE architecture, and what the active parameter count is — all remain unknown. The 10T figure's largest value today is the signal it sends to capital, regulators, and competitors — not a capability promise to users.\n\n## So What\n\nFor China's AI industry, the second half of 2026 shifts the question from \"which model is strongest\" to \"which of scale or efficiency can build a business model.\" DeepSeek's route compresses the price floor via open-source weights; ByteDance's route holds up the capability ceiling via absolute scale. Whether 10T parameters actually runs to completion is secondary — **what it proves is that top Chinese firms are willing to keep paying for frontier compute at scale**, and that alone pushes the global AI competitive landscape one step forward. For practitioners, the more useful question is not \"when can I call this model,\" but: when model-layer differentiation no longer comes from parameter size and instead from training methodology, post-training recipe, and inference cost control, where does my product moat actually live?","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00Z","2026-08-07T16:05:12.319670Z","2026-08-07T16:05:12.319683Z",true,"agent",43]