ByteDance's 10-Trillion-Parameter Bet: Scale Racing Meets Zhang Yiming's "No-Distillation" Stance
On August 7, the Financial Times reported that ByteDance is training an AI model with around 10 trillion parameters, currently in the pre-training stage. Pre-training typically takes 3 to 6 months, followed by fine-tuning and release — meaning this "extra-large model" is still a long way from public availability (Lianhe Zaobao / FT reprint).
The report disclosed several key numbers: at 10T parameters, ByteDance's model is more than three times the size of Moonshot's Kimi K3 (2.8T). Previously, Meituan's LongCat-2.0 and DeepSeek's V4-Pro led the Chinese frontier at 1.6T; Alibaba's latest Qwen3.8-Max sits at around 2.4T (21 Jingji / JRJ). By industry estimates, Anthropic's most advanced Mythos 5 carries around 8T parameters, Fable 5 around 5T — these figures are industry estimates, not publicly confirmed by Anthropic (FourWeekMBA).
What "10 Trillion" Actually Means: Parameter Games in the MoE Era
A necessary clarification: all these numbers refer to total parameters, not "active parameters." In modern frontier sparse MoE architectures, 10T total parameters means the vast majority of weights are "asleep" during any single token forward pass. Total parameter count is closer to a statement — "I can afford this compute, I can stack this much data" — rather than a capability claim.
Put back in the domestic context: DeepSeek V4, Alibaba Qwen3.8-Max, and Moonshot Kimi K3 all pursue sparse MoE + extreme training efficiency — pushing frontier capability with as few active parameters as possible. ByteDance has now chosen the opposite end: stacking total parameters up to 10T. The cost structures diverge sharply — one bets on "capability density per unit inference cost," the other on "absolute capability ceiling."
Zhang Yiming's Day-Before Internal Statement: Refusing Distillation as a Shortcut
Coincidentally, on August 6, The Paper reported Zhang Yiming's rare statement at ByteDance's July all-hands for the Seed team: "ByteDance will not use distillation as a shortcut to lift AI model capability, even if that means we currently lag behind domestic competitors" (Lianhe Zaobao reprint).
Zhang said "making models demands long-termism and delayed gratification — not trading someone else's outputs for a momentary spot on a leaderboard." The report also noted ByteDance internally forbids distillation from open-source models, even enforcing limits via API detection internally — because Zhang believes distillation "interferes with truly long-term technological breakthroughs."
Read together, the 10T bet takes on a clearer annotation: ByteDance chose not to take the shortcut, so it had to commit to the long road. Under pressure from DeepSeek repeatedly raising the open-source capability ceiling and Moonshot's Kimi K3 hitting 57 on the Artificial Analysis Intelligence Index, ByteDance did not take distillation — the most cost-efficient path to capability migration — and instead chose "use absolute scale to substitute for the gap." This road is more expensive, slower, but harder to copy.
The Route Debate: China AI Runs Both Directions at Once
A deeper layer: the truly interesting thing about 10T parameters is not "can China catch up with the US frontier," but "China is internally running two mutually exclusive routes simultaneously."
One is DeepSeek's "small and frugal" thesis — V4 Flash uses 284B total / 13B active MoE, scores 50 on Artificial Analysis Intelligence Index, with blended price around $0.06 per million tokens, ~65% lower single-task cost than GPT-5.6 Luna. The other is ByteDance's "large and fierce" bet — stacking total parameters into Mythos 5's tier (or slightly above by estimates), trading compute density for capability ceiling.
The two routes have different cost curves, target customers, and business models. In the environment where DeepSeek keeps applying price pressure via open-source weights, ByteDance is betting on a different customer class: those willing to pay a premium for "absolute capability ceiling" — Agent workloads, complex reasoning, enterprise-grade RAG. This is why Zhang Yiming's line "willing to sacrifice some short-term returns for long-term goals" only makes sense paired with a 10T compute commitment — without that scale commitment, the statement would just be a pretty phrase.
Compute Sovereignty Lens: Export Controls May Be Thinning
A view that cannot be avoided is compute. The pre-training compute footprint of a 10T-parameter model already approaches the capability band US export controls were designed to constrain. FourWeekMBA flagged this point: if ByteDance is genuinely running MoE pre-training at 10T total parameters, it implies Chinese labs have found ways around dependence on cutting-edge NVIDIA accelerators — whether through stockpiled chips acquired before controls tightened, Huawei Ascend and other domestic substitutes, or training-efficiency improvements that offset the hardware gap.
But several caveats are required: ByteDance has not publicly confirmed the 10T figure, the FT source is "people familiar with the plan"; Mythos 5 and Fable 5 parameter counts are industry estimates, not official disclosures; whether the model will ship at all, whether it remains MoE architecture, and what the active parameter count is — all remain unknown. The 10T figure's largest value today is the signal it sends to capital, regulators, and competitors — not a capability promise to users.
So What
For China's AI industry, the second half of 2026 shifts the question from "which model is strongest" to "which of scale or efficiency can build a business model." DeepSeek's route compresses the price floor via open-source weights; ByteDance's route holds up the capability ceiling via absolute scale. Whether 10T parameters actually runs to completion is secondary — what it proves is that top Chinese firms are willing to keep paying for frontier compute at scale, and that alone pushes the global AI competitive landscape one step forward. For practitioners, the more useful question is not "when can I call this model," but: when model-layer differentiation no longer comes from parameter size and instead from training methodology, post-training recipe, and inference cost control, where does my product moat actually live?