From 'Stacking Hardware' to 'Selling Tokens': China's Compute Industry Is Rewriting Its Billing Logic
If the past three years of China's compute industry were defined by the question of who had the GPUs, the second half of 2026 is rewriting the question. The phrase 'Token factory' — once confined to the internal vocabulary of model labs — has been co-opted by compute providers and is now plastered across their pitch decks.
On the evening of July 29, Xingyun Tech announced that its wholly-owned subsidiary Shenzhen Xingyun had signed a supplementary compute service agreement with VB, a customer described as a leading long-context, foundation-model vendor. The deal expands the contract from 128 to 256 compute units and lifts the 5-year total from approximately 1.0 billion yuan to 3.053 billion yuan — a 201.14% increase. Crucially, the supplement also introduces a new line item: a fixed service fee tied to Token revenue, payable regardless of whether the customer's actual Token business grows or shrinks.
This is not an isolated case. Shanghai Securities News reports that since the start of 2026, more than 10 listed companies — including Runjian, Hongxin Electronics, Xingyun Tech, Chaoxun Technology, and Nanwei Software — have announced plans to explore the 'Token Factory' model. The trend cannot be summarized as 'hardware prices are rising'. The entire billing logic is switching.
Inside the Xingyun Deal: A New 'Token Sharing' Line Item
The supplementary agreement is worth breaking down.
The original contract was a textbook 'sell the resource' deal: 128 compute units × unit price × 5 years = 1.0 billion yuan. The supplement restructures the deal into three tiers:
- Tier 1 — base fixed fee per unit: monthly unit fee rises 21.21% vs. the original contract;
- Tier 2 — expanded unit fee: covering the remaining 120 units plus 128 newly added units, with a 36.36% monthly increase;
- Tier 3 — Token-revenue fixed service fee: payable monthly as long as the customer continues to accept delivery of these compute units, decoupled from actual Token revenue.
On the surface, the decoupling protects Xingyun from the customer's commercial risk. But the definition itself — 'a fee related to the customer's Token revenue' — embeds an option value: if VB's long-text model becomes a hit, Xingyun sits on the upside.
Industry observers read it bluntly: the essence of TaaS (Token-as-a-Service) is that compute companies and model companies share in Token upside — the model company's Token margin flows through to the 'Token factory'. Where the old 'sell-the-resource' model was a one-shot transaction, TaaS aligns the compute provider's economics with the model provider's product-market fit.
Three Layers of a 'Token Factory': Scheduling, Inference Optimization, and Domestic Compute
'Token factory' is not marketing fluff — it is a stack.
Zhang Wen, CEO of Xingyun Tech, recently framed the industry's evolution in three eras: 1.0 was 'selling equipment', 2.0 was 'selling resources', and 3.0 is 'system-integration capabilities decide the winner'. Across the recent announcements, a common architecture emerges:
- Runjian Co.: launched 'Xingsuan Cloud Pool' (星算云池), partnering with major internet platforms; its Wuxiang Yungu intelligent computing center will be upgraded into a 'Token factory';
- Chaoxun Technology: on May 13 signed with Guangxi Nanyi Intelligent to co-build a 'Token factory' covering compute leasing, scheduling, and packaged compute products;
- Hongxin Electronics: building a 'Token factory' in Wuxi, anchored on Huawei super-node compute clusters as first-phase infrastructure;
- Nanwei Software: its Qixingyuan intelligent computing center in Beijing is positioned as a large-scale 'Token factory', intended to anchor a local Token economy.
Stripped down, the shared blueprint is: a domestic/non-US compute substrate at the bottom, a scheduling layer in the middle, and Token-metered, revenue-shared service APIs on top. In the MaaS era all three layers sat inside the model company. In the TaaS era, scheduling and metering move to the compute side; the model company just keeps emitting Tokens.
Why Now: 1000x Token Growth in Two Years, MaaS Gives Way to TaaS
At a forum in April, Chinese Academy of Engineering academician Zheng Weimin offered a key data point: China's Token consumption has grown by three orders of magnitude in two years. That trajectory produces two simultaneous effects:
- Model companies suddenly need long-term, stable compute contracts — not last-minute GPU hunting;
- Compute companies discover that the moment 'Token throughput' becomes a Service Level Agreement (SLA), their pricing power exceeds anything a pure per-GPU-hour contract could deliver.
This is the industry substrate for the TaaS wave in H2 2026 — Token has become a 'unit of electricity' that can be packaged, scheduled, and revenue-shared; compute providers are becoming 'Token power plants'.
One compute-industry source put it more bluntly: many high-end compute service engagements now face delivery shortfalls because most providers cannot secure enough high-performance compute servers. Whoever can secure the hardware commands the pricing power and the long-term contract; downstream large-model players are willing to accept the price hikes to lock in supply. This is a window where hardware scarcity stacks on top of Token demand explosion — and TaaS uses that window to rewrite contract structures.
Will the 'Token Factory' Last?
The risks of the TaaS model are also clear.
First, you are not locking in a customer — you are locking in their hit-making ability. Xingyun's supplement decouples the Token-revenue fee from actual revenue, which looks protective for the compute side, but if VB's long-text model fails to commercialize at scale, renewal and expansion probabilities drop. The 3.05-billion-yuan, 5-year headline number leans heavily on VB's product-market fit.
Second, top-tier foundation model companies may refuse to share Token revenue with compute providers. The leading general-purpose model labs (with proprietary base models and direct distribution channels) have strong Token pricing leverage. They are more likely to self-build or single-source a compute partner than to revenue-share with a 'Token factory'. Today's 'Token factory' customers are mostly specialized long-context / large-context model companies, not the general-purpose frontier.
Third, Token metering and pricing have no industry standard. Different models have wildly different Token lengths, billing rules, and cache hit rates. 'Tokens per watt' is a slick competitive metric, but there is no cross-model, cross-cluster benchmark for it. The industry will inevitably fight over a 'Token unit price' reference standard in the coming year.
So What: TaaS Is Not a New Concept, but It Is Now Written Into Contracts
Step back, and the 'Token factory' is essentially compute moving from a capital good (servers) to a consumer good (API calls).
The IDC and CDN industries walked a similar path — from selling rack space and bandwidth, to billing by QPS and traffic, which spawned Cloudflare, Akamai, and the entire edge-as-a-service category. China's compute industry is now replaying that path in the LLM era: from 'I have H100s / Ascend', to 'I can reliably emit N Tokens per minute'.
Judging from the moves at Xingyun, Runjian, Chaoxun, Hongxin, and Nanwei, China's compute providers have collectively realized that the moat for the next three years will not be GPU count, but scheduling systems and Token metering capability. Over the next 12-18 months, the 'Token factories' that actually run at scale will pull an order-of-magnitude lead over compute companies still operating at the 'I have N cabinets' level.
The single line item — 'a fixed service fee related to the customer's Token revenue' — in Xingyun's contract is the real-world signature of where the era turns.