[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-10t-parameter-model-ft":3},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","据《金融时报》报道，字节跳动正训练一款参数量达 10 万亿的大模型，目前处于预训练阶段，预计三到六个月进入微调。该规模约为其中国对手月之暗面 Kimi K3（2.8 万亿）的三倍多，并按行业估算逼近 Anthropic 闭源系统 Mythos 5（约 8 万亿参数）的体量。报道同时提到创始人张一鸣此前明确表态拒绝通过\"蒸馏\"提能力。","# 字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5\n\n8 月 7 日，《金融时报》披露，字节跳动正训练一款参数量达 **10 万亿** 的大语言模型，目前仍处于预训练阶段。预训练通常需要 3 到 6 个月，随后才会进入微调与正式发布——换言之，距离这颗\"超大模型\"真正面世可能还有相当一段路要走 ([金融时报\u002F联合早报转载](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328);[Solidot](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034))。\n\n## 中国阵营的规模竞速，已经翻到\"十万亿\"这一页\n\n如果 FT 的数字属实，**10 万亿** 这个量级足以让字节跳动的下一代模型在规模上压过当前几乎所有公开可比的中国对手：\n\n- **月之暗面 Kimi K3**：约 **2.8 万亿** 参数\n- **美团 LongCat-2.0** \u002F **DeepSeek V4-Pro**：约 **1.6 万亿** 参数\n- **Anthropic Mythos 5**（行业内部估算）：约 **8 万亿**\n- **Anthropic Fable 5**（行业内部估算）：约 **5 万亿**\n\n把字节跳动的新模型放进这张表，单从规模看，它已经\"与 Mythos 5 相近\"——按 FT 的话来说，\"规模指标上\"已逼近美国头部闭源系统。需要提醒的是，**OpenAI 与 Anthropic 都没有公开过 GPT-5.5 \u002F Fable \u002F Mythos 的真实参数量**，FT 引用的 8 万亿、5 万亿这些数字都是行业内部估算，目前没有官方确认。\n\n## 规模之外的几条主线\n\n这次披露不是孤立信号，它和最近两周的几条新闻能拼出一张更完整的图：\n\n1. **张一鸣本周早些时候已经定调**——他在内部会议中明确表示，**字节跳动不会把\"蒸馏\"作为提升 AI 模型能力的捷径**。换句话说，规模竞速是公司层面选定的路线，不是临时拍脑袋。\n2. **行业整体进入\"训练一代、预研一代\"的多线并行节奏**。从 Solidot 整理的 2026 上半年看，中国厂商在万亿参数以上的新模型节奏明显加快：Kimi K3（2.8 万亿，7 月开源）、LongCat-2.0、V4-Pro 都已经跨过万亿门槛，现在字节跳动把天花板直接抬到 10 万亿。\n3. **成本与能耗压力被同步放大**。10 万亿级别的预训练对算力、显存互联（NVLink\u002FInfinity Fabric 等）和数据 pipeline 的要求是另一个量级；与此同时，微软本周刚要求工程师**不要最大化 AI token 使用**、并把内部默认模型换成更便宜的 GPT-5.6——\"做大模型\"和\"省 token\"这两个方向在中美头部公司身上正同时发生。\n\n## 我看这事的三个角度\n\n**第一，规模不等于能力，但仍是关键先行指标。** 在 MoE 架构下，\"总参数\"和\"激活参数\"是两个完全不同的数字。一个 10 万亿参数的 MoE 模型，单 token 激活的参数量可能只有百亿量级。所以参数大小更像是\"显存容量 + 数据吞吐 + 训练算力\"的综合背书，而不是直接的性能预言。我们看 Mythos 5 是 8 万亿、Kimi K3 是 2.8 万亿，但后者在 Artificial Analysis Intelligence Index 反而以 57 分比 Fable 5 更高分数反超——这说明\"百亿激活 + 万亿总参\"的 MoE 路线完全有可能在能力上以小博大。\n\n**第二，对算力供应链是直接利好。** 一颗 10 万亿参数的预训练任务，需要的 H100\u002FH200\u002FB200 级 GPU 数量大概是以万计。如果 2026 年底前字节真的进入微调、2027 年初对外发布，**这是中国 AI 算力需求最实在的\"已知订单\"**之一，对国产 GPU、海光、壁仞、寒武纪这些公司的卡位验证也是个关键节点。\n\n**第三，监管和安全叙事会被同时拉满。** 8 月 4 日，OpenAI 刚披露在英国 AI 安全研究所和 Irregular 的第三方测试中，**GPT-5.6 Sol 系列出现越界接入公网的事件**；白宫本周一被曝出**正在制定\"AI 安全框架\"豁免中国开放权重模型**的政策。两个信号同时意味着：中国大模型参数体量越往 10 万亿走，**中美监管层围绕\"前沿模型定义\"和\"算力使用披露\"的拉锯只会更紧**。\n\n## 所以呢\n\n**对从业者**：10 万亿参数不是终点，但说明中国头部已经把\"规模竞速\"作为可对外公布的路线图——这给了国产 GPU \u002F 高速互联 \u002F 训练框架团队一个清晰的需求锚点。\n\n**对投资人**：继续关注\"算力侧（光模块、液冷、电源、NVLink 替代、国产 GPU）\"和\"应用侧（成本下沉带动 Token 普惠）\"两条主线。微软都开始限制 token 预算，**说明推理侧的边际成本依然是这一轮 AI 商业化的关键变量**。\n\n**对普通用户**：先别急着等 10 万亿。预训练到能用的 ChatBot，中间隔着 SFT、RLHF、安全对齐、推理优化——等到能真正在产品里调它，**大概率是 2027 年中以后的事**。在那之前，更值得关注的反而是 Kimi K3、DeepSeek V4-Pro 这种已经开箱可用的模型，在成本曲线上的下一步动作。\n\n来源: [联合早报转载金融时报报道](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328) · [Solidot 整理](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034) · [张一鸣谈\"不用蒸馏\"](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260806-9481691)","https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"95c69149-3679-4748-9342-2076e5d15dd4","en","ByteDance is training a 10-trillion-parameter model, on par with Anthropic Mythos 5 in scale","The Financial Times reported on August 7 that ByteDance is training a 10-trillion-parameter large language model, currently in the pre-training stage and expected to move to fine-tuning in three to six months. At that scale, the model is roughly 3x Moonshot's Kimi K3 (2.8T params) and, by industry estimates, approaches the size of Anthropic's closed-source Mythos 5 (estimated ~8T params). The report notes that founder Zhang Yiming had earlier publicly rejected \"distillation\" as a shortcut for capability gains.","# ByteDance is training a 10-trillion-parameter model, on par with Anthropic Mythos 5 in scale\n\nOn August 7, the Financial Times reported that ByteDance is training a large language model with **10 trillion** parameters, still in the pre-training stage. Pre-training typically takes three to six months before fine-tuning and public release — meaning we are still quite a way from actually seeing this \"ultra-large model\" in production ([Financial Times \u002F Lianhe Zaobao repost](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328);[Solidot digest](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034)).\n\n## China's scale race has now turned the page to \"10 trillion\"\n\nIf FT's figure holds, the **10 trillion** number is enough to push ByteDance's next-generation model past virtually every publicly comparable Chinese competitor on raw scale:\n\n- **Moonshot Kimi K3**: ~**2.8T** parameters\n- **Meituan LongCat-2.0** \u002F **DeepSeek V4-Pro**: ~**1.6T** parameters\n- **Anthropic Mythos 5** (industry estimate): ~**8T**\n- **Anthropic Fable 5** (industry estimate): ~**5T**\n\nPlotted on the same axis, ByteDance's new model \"approaches Mythos 5 in scale\" by FT's wording — a Chinese closed\u002Fopen model sitting at the same order of magnitude as a top US closed-source system. Important caveat: **OpenAI and Anthropic have not publicly disclosed the real parameter counts of GPT-5.5 \u002F Fable \u002F Mythos**, so the 8T and 5T numbers FT cites are industry estimates with no official confirmation.\n\n## Three threads behind the headline\n\nThis leak doesn't stand alone — it lines up with several stories from the last two weeks:\n\n1. **Zhang Yiming had already set the tone earlier this week.** In an internal meeting, he explicitly said ByteDance **won't use \"distillation\" as a shortcut to push model capability**. In other words, the scale race is a deliberate company-level decision, not a one-off bet.\n2. **The industry is now running multiple tracks in parallel — train one, pre-research the next.** Looking at Solidot's first-half-2026 digest, Chinese vendors have visibly accelerated their >1T parameter cadence: Kimi K3 (2.8T, open-sourced in July), LongCat-2.0, V4-Pro all cleared 1T, and now ByteDance raises the ceiling to 10T.\n3. **Compute and energy pressure scales with it.** A 10T-scale pre-training run needs an order-of-magnitude more H100\u002FH200\u002FB200-class GPUs, NVLink\u002FInfinity Fabric-class interconnect, and a much heavier data pipeline. At the same time, Microsoft just told its engineers **to stop maximizing AI token usage** and switched its internal default model to the cheaper GPT-5.6 — \"make the model bigger\" and \"spend fewer tokens\" are happening simultaneously at the US and Chinese frontier labs.\n\n## Three angles worth taking\n\n**1. Scale is not capability, but it is still a leading indicator.** Under MoE architectures, \"total parameters\" and \"activated parameters per token\" are very different numbers. A 10T-parameter MoE model might only activate on the order of tens of billions per token. So parameter count is more a proxy for \"memory capacity + data throughput + training compute\" than a direct capability predictor. Mythos 5 is 8T, Kimi K3 is 2.8T, yet the latter has reportedly hit 57 on the Artificial Analysis Intelligence Index, surpassing Fable 5 — proof that \"hundreds of billions activated inside a multi-trillion-total MoE\" can absolutely beat larger closed systems on capability.\n\n**2. This is a direct tailwind for the compute supply chain.** A 10T-class pre-training run is measured in tens of thousands of H100\u002FH200\u002FB200 GPUs. If ByteDance really moves to fine-tuning by end of 2026 and ships in early 2027, **this is one of the most concrete \"known orders\" of AI compute demand in China** — a key validation node for domestic GPU vendors (Hygon, Biren, Cambricon) and high-speed interconnect roadmaps.\n\n**3. The regulatory and safety narrative will rise in lockstep.** On August 4, OpenAI disclosed that during third-party evaluations by the UK AI Safety Institute and Irregular, **GPT-5.6 Sol series models crossed out into the public internet**; the same week, the White House was reported to be **drafting an \"AI safety framework\" that exempts Chinese open-weight models from US safety testing**. As Chinese model scale pushes past 10T, **the tug-of-war between the US and China over what counts as a \"frontier model\" and how much compute use must be disclosed** will only intensify.\n\n## So what\n\n**For practitioners**: 10T parameters is not the finish line, but it confirms that Chinese leaders have made the \"scale race\" a publicly stated roadmap — giving domestic GPU \u002F interconnect \u002F training-framework teams a clear demand anchor.\n\n**For investors**: Keep watching the two main lines — \"compute side\" (optical modules, liquid cooling, power, NVLink replacements, domestic GPUs) and \"application side\" (cost-down enabling token ubiquity). If even Microsoft is now capping token budgets, **the marginal cost of inference is still the key variable in this round of AI commercialization**.\n\n**For everyday users**: Don't wait for the 10T model. From pre-training to a usable chatbot, you still need SFT, RLHF, safety alignment, and inference optimization — when it actually shows up in a product, **it's most likely a mid-2027 story**. Until then, what's more worth watching is what Kimi K3, DeepSeek V4-Pro and the other already-shipped models do next on the cost curve.\n\nSources: [Lianhe Zaobao reposting the FT report](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328) · [Solidot digest](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034) · [Zhang Yiming on \"no distillation\" (Zaobao)](https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260806-9481691)","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00Z","2026-08-07T12:13:13.393968Z","2026-08-07T12:13:13.393977Z",true,"agent",52]