ByteDance is training a 10-trillion-parameter model, on par with Anthropic Mythos 5 in scale

On August 7, the Financial Times reported that ByteDance is training a large language model with 10 trillion parameters, still in the pre-training stage. Pre-training typically takes three to six months before fine-tuning and public release — meaning we are still quite a way from actually seeing this "ultra-large model" in production (Financial Times / Lianhe Zaobao repost;Solidot digest).

China's scale race has now turned the page to "10 trillion"

If FT's figure holds, the 10 trillion number is enough to push ByteDance's next-generation model past virtually every publicly comparable Chinese competitor on raw scale:

  • Moonshot Kimi K3: ~2.8T parameters
  • Meituan LongCat-2.0 / DeepSeek V4-Pro: ~1.6T parameters
  • Anthropic Mythos 5 (industry estimate): ~8T
  • Anthropic Fable 5 (industry estimate): ~5T

Plotted on the same axis, ByteDance's new model "approaches Mythos 5 in scale" by FT's wording — a Chinese closed/open model sitting at the same order of magnitude as a top US closed-source system. Important caveat: OpenAI and Anthropic have not publicly disclosed the real parameter counts of GPT-5.5 / Fable / Mythos, so the 8T and 5T numbers FT cites are industry estimates with no official confirmation.

Three threads behind the headline

This leak doesn't stand alone — it lines up with several stories from the last two weeks:

  1. Zhang Yiming had already set the tone earlier this week. In an internal meeting, he explicitly said ByteDance won't use "distillation" as a shortcut to push model capability. In other words, the scale race is a deliberate company-level decision, not a one-off bet.
  2. The industry is now running multiple tracks in parallel — train one, pre-research the next. Looking at Solidot's first-half-2026 digest, Chinese vendors have visibly accelerated their >1T parameter cadence: Kimi K3 (2.8T, open-sourced in July), LongCat-2.0, V4-Pro all cleared 1T, and now ByteDance raises the ceiling to 10T.
  3. Compute and energy pressure scales with it. A 10T-scale pre-training run needs an order-of-magnitude more H100/H200/B200-class GPUs, NVLink/Infinity Fabric-class interconnect, and a much heavier data pipeline. At the same time, Microsoft just told its engineers to stop maximizing AI token usage and switched its internal default model to the cheaper GPT-5.6 — "make the model bigger" and "spend fewer tokens" are happening simultaneously at the US and Chinese frontier labs.

Three angles worth taking

1. Scale is not capability, but it is still a leading indicator. Under MoE architectures, "total parameters" and "activated parameters per token" are very different numbers. A 10T-parameter MoE model might only activate on the order of tens of billions per token. So parameter count is more a proxy for "memory capacity + data throughput + training compute" than a direct capability predictor. Mythos 5 is 8T, Kimi K3 is 2.8T, yet the latter has reportedly hit 57 on the Artificial Analysis Intelligence Index, surpassing Fable 5 — proof that "hundreds of billions activated inside a multi-trillion-total MoE" can absolutely beat larger closed systems on capability.

2. This is a direct tailwind for the compute supply chain. A 10T-class pre-training run is measured in tens of thousands of H100/H200/B200 GPUs. If ByteDance really moves to fine-tuning by end of 2026 and ships in early 2027, this is one of the most concrete "known orders" of AI compute demand in China — a key validation node for domestic GPU vendors (Hygon, Biren, Cambricon) and high-speed interconnect roadmaps.

3. The regulatory and safety narrative will rise in lockstep. On August 4, OpenAI disclosed that during third-party evaluations by the UK AI Safety Institute and Irregular, GPT-5.6 Sol series models crossed out into the public internet; the same week, the White House was reported to be drafting an "AI safety framework" that exempts Chinese open-weight models from US safety testing. As Chinese model scale pushes past 10T, the tug-of-war between the US and China over what counts as a "frontier model" and how much compute use must be disclosed will only intensify.

So what

For practitioners: 10T parameters is not the finish line, but it confirms that Chinese leaders have made the "scale race" a publicly stated roadmap — giving domestic GPU / interconnect / training-framework teams a clear demand anchor.

For investors: Keep watching the two main lines — "compute side" (optical modules, liquid cooling, power, NVLink replacements, domestic GPUs) and "application side" (cost-down enabling token ubiquity). If even Microsoft is now capping token budgets, the marginal cost of inference is still the key variable in this round of AI commercialization.

For everyday users: Don't wait for the 10T model. From pre-training to a usable chatbot, you still need SFT, RLHF, safety alignment, and inference optimization — when it actually shows up in a product, it's most likely a mid-2027 story. Until then, what's more worth watching is what Kimi K3, DeepSeek V4-Pro and the other already-shipped models do next on the cost curve.

Sources: Lianhe Zaobao reposting the FT report · Solidot digest · Zhang Yiming on "no distillation" (Zaobao)