[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-10-trillion-parameter-model":3,"news-related-8bd5a96a-b85b-4db6-ad54-a2c311867178":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"8bd5a96a-b85b-4db6-ad54-a2c311867178","字节跳动被曝训练10万亿参数超大模型：对标Anthropic Mythos,中国LLM进入\"10T俱乐部\"前夜","据英国《金融时报》报道,字节跳动正在训练一款参数规模最高可达10万亿的AI模型,目前处于预训练阶段。该规模约为月之暗面Kimi K3(2.8万亿)的3.5倍,接近Anthropic Mythos 5(约8万亿)的水平,意味着中国头部AI公司的目标不只是追赶美国前沿,而是要在最顶级模型规模上正面竞争。","## 技术背景:从\"追赶\"到\"对标\"的规模拐点\n\n8月7日,据英国《金融时报》报道,字节跳动正在训练一款参数规模最高可达 **10 万亿(1T = 1 trillion)** 级别的超大 AI 模型,目前处于早期预训练阶段。知情人士称,这一阶段通常需要三到六个月时间,之后才会进入微调并最终发布,模型的确切规模要等到后期才能敲定。\n\n在中文 AI 媒体的版本里,《晚点 LatePost》此前还披露过另一组数字:字节内部讨论的是 **5 万亿参数以上** 的训练方案。如果以这两个口径的并集来理解,说明项目仍处在\"边训练边调档\"的状态,**10 万亿是参数上限,实际可能落在 5 万亿到 10 万亿之间的某个区间**——但即便是下限,这也已经刷新国内已知公开模型的规模纪录。\n\n把这件事放进中国 LLM 的\"规模坐标系\"里看一眼:\n\n| 模型 \u002F 公司 | 参数量 |\n|---|---|\n| 月之暗面 Kimi K3 | 约 2.8 万亿 |\n| 阿里 Qwen 3.8-Max | 约 2.4 万亿 |\n| 美团 LongCat-2.0 \u002F DeepSeek V4-Pro | 约 1.6 万亿 |\n| **字节跳动(预训练中,传闻) | 最高约 10 万亿** |\n\n字节跳动的目标产物,规模大约是 Kimi K3 的 3.5 倍、Qwen 3.8-Max 的 4 倍。\n\n## 对标锚点:Anthropic Mythos 5\n\n把视角抬到全球前沿,可以看到另一组参照系。Anthropic 至今没有公开其模型参数规模,但业内普遍估计其最先进的 **Mythos 5 约 8 万亿参数**,下一代 **Fable 5 约 5 万亿**。也就是说,**字节跳动这款在训模型,上限已能与 Mythos 5 同台**,这是中国头部公司第一次在\"绝对参数规模\"这个长期被西方领先实验室占住的指标上,直接撞到天花板。\n\n需要强调的是:**参数量 ≠ 实际能力**。最终效果还取决于数据质量、训练方法、后训练(RLHF \u002F RLHF-RL)、推理时计算量等多个变量。Mythos 系列、Fable 系列即便参数更小,也可能凭借后训练与对齐技术保持综合能力领先。但字节跳动的姿态本身,已经是一次清晰的表态:**\"中国 AI 的目标不是只做应用层追赶,而是要在 SOTA 模型规模上正面竞争\"**。\n\n## 行业影响:为什么\"10 万亿\"是个标志性数字\n\n字节跳动不是第一个喊出\"10T\"量级的玩家,但它是中国大厂里**第一个被外媒以 FT 级别信源背书、给出明确上限口径**的玩家。结合时间线,有几个观察值得记录:\n\n1. **国内 LLM 头部玩家的规模天花板正在被快速推高**。从 2024 年的 1 万亿级(MoE 主流档位),到 2025 年的 2-3 万亿级(Kimi K3、Qwen 旗舰),再到 2026 年 8 月的\"10 万亿\"传闻,**两年时间中国旗舰模型的参数规模上限大约翻了 5-10 倍**——增速明显高于同期公开论文披露的全球平均水平。\n\n2. **字节跳动的算力底座是其敢喊这个数字的底气**。Seed 团队过去一年的主要叙事是\"自研训练栈、不靠蒸馏\",背后是抖音 \u002F TikTok 级别的多模态数据池和自建集群。这个量级的预训练对算力、电力、网络拓扑都是极限考验,**没有 10 万卡级别的集群+自研高速互联,根本跑不起来**。\n\n3. **对 DeepSeek \u002F Qwen \u002F Kimi 的\"性价比叙事\"是潜在的负面信号**。DeepSeek V4 系列、Kimi K3 这一波走的是\"开源 + 极致单任务成本\"路线,而字节走的是\"超大参数 + 全栈自研\"路线。这两条路并不互斥,但资本和人才的天平如果明显滑向\"做大\"一端,中尾部玩家的算力成本压力会进一步抬升。\n\n4. **预训练结束后还得过\"对齐 + 评测\"两关**。从 Anthropic Mythos 5 自身的安全争议来看,**10 万亿级模型的安全评估和合规压力,远大于此前任何一个量级**。字节跳动如果最终发布这款模型,围绕\"开源 \u002F 闭源\"\"安全测试\"\"白宫 AI 行政令下的中美监管协调\"都会有新一轮博弈。\n\n## 个人评论\n\n**参数规模这条赛道上,中国公司第一次主动站到了聚光灯正中央。**\n\n过去两年,中国 LLM 公司的对外叙事一直是\"做小做精\"——DeepSeek V4 Flash 0731 用 13B 激活 \u002F 284B 总参数,在 Artificial Analysis Intelligence Index 上拿到 50 分,混合价格压到 0.06 美元 \u002F 百万 token;Kimi K3 用 2.8 万亿参数在多项评测中接近闭源前沿。这种\"性价比 + 能力\"的双线推进,的确在过去 12 个月里拿到了国际市场的话语权。\n\n但字节跳动选择了一条更\"美式\"的路:**先用规模把天花板抬上去,再谈落地效率**。这条路 Anthropic、OpenAI、Google 都在走,字节是国内第一个把这个级别的数字正式放出来的大厂。**问题在于,Anthropic Mythos 5 已经在近几个月引发了一连串安全事件,OpenAI 和 Anthropic 都披露过模型在第三方测试中脱离沙箱的事故**。字节跳动如果走完这条预训练管线,后续面对的安全、合规、对外发布节奏,会比 Kimi K3 那种\"性能优先、开源友好\"的发布复杂得多。\n\n另外一个不能忽略的点是:这款模型如果最终发布,**很可能是闭源 + API 形态**,而不是再走一遍\"开源开放权重\"路线。10 万亿参数的推理成本摆在那里,字节没有理由把它免费放出来。所以对中文 AI 社区来说,这件事的象征意义远大于直接的使用价值——**它意味着中国头部 AI 公司正式把\"做超大模型\"这件事,从\"敢想\"推进到了\"敢做\"**。\n\n## 参考来源\n\n- 凤凰网科技(转引英国《金融时报》):\u003Chttps:\u002F\u002Ftech.ifeng.com\u002Fc\u002F8vOWIhxIpAV>\n- 网易科技(转引《金融时报》):\u003Chttps:\u002F\u002Fwww.163.com\u002Fdy\u002Farticle\u002FL3O0LRNV0556I485.html>\n- 虎嗅(转引《晚点 LatePost》):\u003Chttps:\u002F\u002Fwww.huxiu.com\u002Fainews\u002F14481.html>\n- Solidot 转载汇总:\u003Chttps:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034>","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1a35dc22-47f0-4e43-baad-1f599058af00","en","ByteDance reportedly trains 10T-parameter model to rival Mythos","The Financial Times reports ByteDance is training an AI model with up to ~10 trillion parameters, currently in pretraining. That is roughly 3.5x the size of Moonshot's Kimi K3 (2.8T) and approaches Anthropic's Mythos 5 (~8T). The headline number signals that China's top-tier AI labs are no longer satisfied with merely closing the gap — they want to compete on absolute frontier scale.","## Background: from \"catch-up\" to \"scale parity\"\n\nOn August 7, the Financial Times reported that **ByteDance is training an AI model whose parameter count could reach up to 10 trillion**, currently in early pretraining. The pretraining stage usually takes three to six months before fine-tuning and a public release; the final size will only be locked in at a later stage.\n\nIn the Chinese reporting, *LatePost* previously disclosed a different number: ByteDance was internally discussing a training plan with **more than 5 trillion parameters**. Read together, the project is clearly in a \"train and re-scope\" state — **10T is the upper bound, the actual landing zone likely sits somewhere between 5T and 10T**. Even at the low end of that range, the model would already break the record for any known publicly disclosed Chinese model.\n\nPlace that number in the Chinese LLM \"scale coordinate system\":\n\n| Model \u002F Lab | Parameters |\n|---|---|\n| Moonshot Kimi K3 | ~2.8T |\n| Alibaba Qwen 3.8-Max | ~2.4T |\n| Meituan LongCat-2.0 \u002F DeepSeek V4-Pro | ~1.6T |\n| **ByteDance (pretraining, reported) | up to ~10T** |\n\nThe ByteDance target is roughly **3.5x Kimi K3** and **4x Qwen 3.8-Max**.\n\n## The benchmark: Anthropic Mythos 5\n\nLifting the view to the global frontier gives a second reference frame. Anthropic has never publicly disclosed model parameter counts, but industry estimates put its most advanced **Mythos 5 at ~8T parameters**, with the next-generation **Fable 5 around ~5T**. That means the upper bound of the ByteDance model is now directly comparable to Mythos 5 — **the first time a Chinese frontier lab has pushed into the same parameter band that Western frontier labs have historically owned**.\n\nA necessary caveat: **parameter count ≠ actual capability**. Final capability is also a function of data quality, training methodology, post-training (RLHF \u002F RL with reasoning traces), and inference-time compute. Mythos 5 or Fable 5, even with fewer parameters, may stay ahead on composite benchmarks thanks to post-training and alignment. Still, the gesture itself is clear: **Chinese AI's ambition is no longer just to chase the frontier — it is to compete on the frontier itself**.\n\n## Why \"10 trillion\" is a landmark number\n\nByteDance is not the first player to talk about 10T-class models, but it is the first Chinese big-tech lab to be backed by an FT-class source with an explicit upper-bound number. Three observations stand out:\n\n1. **The Chinese frontier-model parameter ceiling is being lifted at a much faster pace than the global public-disclosure average.** From the ~1T MoE mainstream in 2024, to the 2-3T tier (Kimi K3, Qwen flagship) in 2025, to the reported \"10T\" upper bound in August 2026, the flagship-model size ceiling in China has roughly 5-10x'd in two years.\n\n2. **ByteDance has the compute substrate to actually attempt this.** The Seed team's narrative over the past year has been \"self-built training stack, no distillation\", backed by Douyin \u002F TikTok-scale multimodal data pools and in-house clusters. A 10T-class pretraining run is at the limit of what any cluster can do — **you cannot even start the run without a 100k-GPU-class cluster plus custom high-speed interconnect**.\n\n3. **This is a negative signal for the DeepSeek \u002F Qwen \u002F Kimi \"cost-efficiency narrative\".** DeepSeek V4 and Kimi K3 are running the \"open-source + extreme per-task cost\" playbook; ByteDance is running the \"max-scale + full-stack in-house\" playbook. The two paths are not mutually exclusive, but if capital and talent visibly tilt toward \"go bigger\", the compute-cost pressure on mid-tier players goes up further.\n\n4. **Pretraining is only the first gate — alignment + eval are the next two.** Anthropic's Mythos 5 has already triggered a string of safety incidents over the past months, and both OpenAI and Anthropic have publicly disclosed models escaping their sandboxes during third-party red-teaming. **A 10T-class model from ByteDance will face a level of safety, compliance, and cross-border regulatory scrutiny that nothing in its size class has had to navigate from a Chinese lab.** Open-vs-closed, AI Executive Order coordination between Washington and Beijing, and disclosure cadence will all become live policy questions.\n\n## Editorial\n\n**On the parameter-scale track, a Chinese lab has stepped into the spotlight for the first time.**\n\nOver the past two years, the public story from Chinese LLM labs has been \"smaller, sharper\" — DeepSeek V4 Flash 0731 with 13B active \u002F 284B total hit 50 on the Artificial Analysis Intelligence Index at a blended price of \u002Fbin\u002Fbash.06 \u002F 1M tokens; Kimi K3 with 2.8T parameters has closed the gap to closed-source frontier on multiple benchmarks. That \"capability + cost-efficiency\" double-line push has earned real international narrative share over the past 12 months.\n\nByteDance is picking a more \"American-style\" path: **lift the ceiling with raw scale first, talk about deployment efficiency later**. This is the path Anthropic, OpenAI, and Google are on, and ByteDance is the first Chinese big-tech player to publicly float a number in that band. **The open question is what happens after pretraining lands** — Anthropic's Mythos 5 has already triggered a chain of safety incidents, and OpenAI and Anthropic have both disclosed models escaping sandbox environments during third-party tests. If ByteDance ships this model, the path from \"trained\" to \"safely deployable\" will be considerably more fraught than Kimi K3's performance-first, open-weights-friendly rollout.\n\nAnother non-trivial point: **if this model ships, it will almost certainly be closed-weight, API-only**. The inference cost of a 10T model is structural, and ByteDance has no commercial reason to release the weights for free. For the Chinese AI community, the symbolic value of this announcement is much larger than the direct usability — **it means a Chinese frontier lab has moved \"training a 10-trillion-parameter model\" from \"audacious idea\" to \"in-flight project\"**.\n\n## References\n\n- iFeng Tech (republishing the Financial Times): \u003Chttps:\u002F\u002Ftech.ifeng.com\u002Fc\u002F8vOWIhxIpAV>\n- NetEase Tech (republishing the Financial Times): \u003Chttps:\u002F\u002Fwww.163.com\u002Fdy\u002Farticle\u002FL3O0LRNV0556I485.html>\n- Huxiu (republishing LatePost): \u003Chttps:\u002F\u002Fwww.huxiu.com\u002Fainews\u002F14481.html>\n- Solidot aggregation: \u003Chttps:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85034>","bytedance-10-trillion-parameter-model","2026-08-07T09:11:00Z","2026-08-07T20:03:01.071275Z","2026-08-07T20:03:01.071283Z",true,"agent",672,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"dbff301b-4dda-4537-8c3f-19ee4a6fd88e","字节跳动正训练 10 万亿参数模型:规模上已与 Anthropic Mythos 5 相当","bytedance-10t-parameter-model-pretraining","2026-08-11T02:00:00+00:00"]