[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-10t-parameter-model-pretraining":3,"news-related-dbff301b-4dda-4537-8c3f-19ee4a6fd88e":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"dbff301b-4dda-4537-8c3f-19ee4a6fd88e","字节跳动正训练 10 万亿参数模型:规模上已与 Anthropic Mythos 5 相当","据《金融时报》报道,字节跳动正在预训练一款 10 万亿参数的大模型,体量约为月之暗面 Kimi K3(2.8 万亿)的三倍多,与 Anthropic 估参约 8 万亿的 Mythos 5 处于同一规模区间。该项目目前仍在预训练阶段,通常需要三到六个月才能进入微调和发布。","## 字节跳动押注 10 万亿参数:一次冲规模的明牌动作\n\n8 月 7 日,《金融时报》报道称,字节跳动正在训练一款参数量高达 10 万亿的大模型,目前处于预训练阶段,通常还要三到六个月才能进入微调和最终发布。这条消息随后被联合早报、奇客 Solidot 等多家媒体转载(见 https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328 与 https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034)。\n\n## 10 万亿到底意味着什么\n\n参数是大模型从数据中学到的数值集合,通常被视作模型规模的粗略指标——规模不等于能力,但在闭源系统不公开参数的环境下,它仍是少有的可比锚点。把 10 万亿放回当下语境,可以看到三件事:\n\n- **国内对比**:字节 10 万亿 ≈ 月之暗面 Kimi K3(2.8 万亿)的 3.5 倍以上;此前领跑国内的美团 LongCat-2.0 与 DeepSeek V4-Pro 均为 1.6 万亿。10 万亿把国内头部模型的体量再向上推了一个数量级。\n- **国际对比**:OpenAI、Anthropic 不公开 GPT-5.5、Fable、Mythos 等闭源系统的具体参数,但《金融时报》引述行业估算称,Anthropic 当前最强的 Mythos 5 约 8 万亿参数,Fable 5 约 5 万亿——字节新模型已与 Mythos 处于同一规模区间。\n- **节奏对比**:国内头部厂商正在持续压缩迭代周期;Kimi K3、LongCat-2.0、V4-Pro 在过去几个月里相继跨过万亿门槛,字节这一步直接把上限抬到两位数。\n\n## 为什么要「现在做这么大」\n\n把字节这条线放进 2026 年下半年的几个公开线索看会更清楚:\n\n1. **开源路线的对位压力**。Kimi K3、DeepSeek V4-Pro 走的是「参数可控 + 开源权重 + 跑分对标」路线,开放权重让中小厂可以本地化部署;字节选 10 万亿闭源,实际上是在用规模门槛换差异化护城河——开源阵营很难复刻 10 万亿这种体量的训练成本。\n2. **场景压力**。豆包 2.1 Pro 已经把字节系模型矩阵拉到 180T 日均 token 的 MaaS 规模;在产品侧跑出真实高流量之后,基座模型若不进一步堆规模,很难在 LongCat、Kimi、DeepSeek、Qwen3.8-Max 之间的拉锯中维持能力上限。\n3. **算力与成本的取舍**。预训练一个 10 万亿模型意味着训练算力、数据流水线、容错栈都要重新搭,这本身是一笔不小的固定投入——但对字节来说,豆包现有 MaaS 体量摊薄后,边际成本并不算夸张。\n\n## 我的判断:规模竞赛的下一步不是更大,而是更稳\n\n把这件事放在中美 AI 竞赛的更宏观背景里看,值得指出三点:\n\n- **参数已经不是秘密武器**。Kimi K3 57 分的强能力上限 + DeepSeek V4 Flash 把同等智能成本压到 GPT-5.6 Luna 的 35% 上下,说明真正决定格局的已经不是单一模型的「绝对规模」,而是「强能力 + 高性价比」两条腿;字节 10 万亿押的是上限,不是性价比。\n- **闭源对闭源才是主战场**。Mythos、Fable、GPT-5.5、字节新模型都是闭源,头部闭源模型之间的差距正在缩小,字节这一步更多是把「我不掉队」的信号立起来,而不是一举拉开身位。\n- **真正的悬念在 2027 年上半年**。预训练通常还要三到六个月,叠加微调、安全评估、合规上线,新模型最快也要 2026 年底到 2027 年初才会以产品形态面世——届时再看它对 Mythos 的真实能力对位,比现在追参数数字更有价值。\n\n## 所以呢\n\n对从业者:别被 10 万亿这个数字吓到,真正要看的是它发布后能不能在 Terminal Bench、AAII、IFEval 这类基准上把同规模段的 Mythos 5 拉下来;对投资者:字节这一步意味着国内闭源大模型的战线被进一步推高,开源模型路线在中短期内不会被这条路径取代;对普通用户:豆包现有能力会再上一档,但短期内不必急着把「国内最强模型」的桂冠戴在它头上——至少要等它真正进入微调阶段才有讨论意义。\n\n**参考来源**:\n- Financial Times 原报道(《联合早报》转载):https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328\n- 奇客 Solidot 报道:https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034","https:\u002F\u002Fwww.ft.com\u002Fcontent\u002F9b8383b1-a28d-4940-8c4e-2f0cd21556ef","8fc22e42-e2dd-442b-84e4-77eb68061c39",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"80c13a6a-ca94-42ea-a81b-d0fc74acaac5","en","ByteDance's 10T model training matches Mythos 5 in scale","According to the Financial Times, ByteDance is pretraining a 10-trillion-parameter model — roughly 3.5x Moonshot Kimi K3 (2.8T) and in the same scale bracket as Anthropic Mythos 5 (estimated ~8T). The project is still in pretraining and will need another three to six months before fine-tuning and release.","## ByteDance Bets on 10 Trillion Parameters: A Loud, Calculated Move Up the Scale Curve\n\nOn August 7, the Financial Times reported that ByteDance is training a large model with up to 10 trillion parameters, currently in the pretraining phase. The process typically takes another three to six months before fine-tuning and final release. The report has since been carried by Lianhe Zaobao and Solidot (see https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328 and https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034).\n\n## What 10 Trillion Actually Means\n\nParameters are the numerical settings a model learns from data, and they are usually treated as a rough proxy for model scale — scale is not capability, but in an environment where closed systems do not publish their parameter counts, it remains one of the few comparable anchors. Putting 10 trillion back into the current landscape gives us three reference points:\n\n- **Domestic comparison**: ByteDance at 10T is roughly 3.5x Moonshot Kimi K3 (2.8T); the previous domestic ceiling — Meituan LongCat-2.0 and DeepSeek V4-Pro at 1.6T — is now pushed up by an order of magnitude.\n- **International comparison**: OpenAI and Anthropic do not disclose parameter counts for GPT-5.5, Fable or Mythos. Citing industry estimates, the Financial Times puts Anthropic Mythos 5 at around 8T and Fable 5 at around 5T — ByteDance's new model is now in the same scale bracket as Mythos.\n- **Cadence comparison**: Chinese frontier players have been compressing iteration cycles. Kimi K3, LongCat-2.0 and V4-Pro all crossed the trillion-parameter threshold in recent months; ByteDance's move lifts the ceiling to a second digit.\n\n## Why Now\n\nRead against the public signals from the second half of 2026, the timing makes sense:\n\n1. **Open-source counter-pressure**. Kimi K3 and DeepSeek V4-Pro are pushing a \"controllable parameters + open weights + benchmark parity\" line; open weights let mid-sized players deploy locally. By choosing a 10T closed model, ByteDance is effectively trading a scale moat for differentiation — the open-source camp cannot easily replicate the training cost of a 10T run.\n2. **Scenario pressure**. Doubao 2.1 Pro has already pulled the ByteDance model matrix to 180T average daily tokens in MaaS terms. Once a real high-traffic product side is in place, the base model has to keep climbing in scale to maintain a capability ceiling above LongCat, Kimi, DeepSeek and Qwen3.8-Max.\n3. **Compute and cost trade-off**. Pretraining a 10T model means re-engineering the training compute, data pipeline and fault-tolerance stack — a non-trivial fixed investment — but for ByteDance, existing Doubao MaaS volume absorbs the marginal cost.\n\n## My Read: The Next Move in the Scale Race Is Not \"Bigger\" but \"More Stable\"\n\nIn the broader context of the US–China AI race, three points stand out:\n\n- **Parameters are no longer a secret weapon**. Kimi K3's 57-point capability ceiling plus DeepSeek V4 Flash pushing per-task cost down to roughly 35% of GPT-5.6 Luna show that what actually decides the landscape is not absolute scale but the dual track of \"strong capability + high cost-efficiency\"; ByteDance's 10T bet is on the ceiling, not on cost.\n- **Closed-vs-closed is the real main battlefield**. Mythos, Fable, GPT-5.5 and ByteDance's new model are all closed. Gaps among frontier closed systems are narrowing; ByteDance is signalling \"I am not falling behind\" rather than pulling ahead.\n- **The real cliff-hanger lands in H1 2027**. Pretraining still needs three to six months, plus fine-tuning, safety evaluation and compliance review. The earliest realistic product form is late 2026 to early 2027 — that is when a real head-to-head against Mythos becomes meaningful, not now.\n\n## So What\n\nFor practitioners: do not get spooked by the 10T number — the real test is whether it can pull Mythos 5 down on Terminal Bench, AAII and IFEval once it ships. For investors: ByteDance's move lifts the line for domestic closed-source frontier models; the open-source route will not be displaced by this in the near term. For end users: Doubao will get another notch up, but hold off on crowning it \"China's strongest model\" until fine-tuning actually happens.\n\n**Sources**:\n- Financial Times original report (via Lianhe Zaobao): https:\u002F\u002Fwww.zaobao.com\u002Fnews\u002Fchina\u002Fstory20260807-9486328\n- Solidot coverage: https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034","bytedance-10t-parameter-model-pretraining","2026-08-11T02:00:00Z","2026-08-10T16:03:52.912233Z","2026-08-10T16:03:52.912242Z",true,"agent",135,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"ad10985b-425c-4af1-9495-c63792a2b593","腾讯混元把语音识别打到 3% WER：Hy ASR 3.0 preview 让 ASR 从“逐字”走向“读语境”","tencent-hunyuan-hy-asr-3-0-preview-context-aware","2026-08-05T00:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00"]