On May 20, at the 2026 Alibaba Cloud Summit, Alibaba released a 128-card super-node server based on the next-gen Pingtouge AI chip Zhenwu M890. The server carries the interconnect chip ICN Switch 1.0, with communication latency as low as the hundred-nanosecond range, and can compose 128 AI chips into a single unified-scheduling compute cluster — designed to meet the dual demands of massive concurrent inference and LLM training in the agentic era.
Traditional LLM training and inference often face communication bandwidth bottlenecks at the chip level — when cross-node collaboration becomes the norm, inter-machine communication latency directly offsets the gains of compute scaling. Zhenwu M890 achieves hundred-nanosecond-level latency at the interconnect layer via the in-house ICN Switch 1.0, meaning data exchange among 128 cards can proceed with virtually no wait. This is especially critical for Transformer models that need to frequently transfer attention weights or KV Cache across nodes.
From a system-architecture perspective, Alibaba Cloud's 128-card super-node thinking echoes NVIDIA's recently delivered Agent-dedicated CPU Vera: NVIDIA rethinks concurrent scheduling for agent scenarios from the processor layer, while Alibaba solves the communication dead spots of multi-chip collaboration from the interconnect layer. Both are tackling the same fundamental problem — when AI workloads shift from single-model inference to multi-agent concurrency, the existing infrastructure assumptions are no longer enough.
Notably, this is the first product-scale landing of Pingtouge chips at the super-node level. From Yitian CPUs to Han Guang NPUs to Zhenwu M890, Pingtouge's chip iteration path is moving from single-chip performance optimization to system-level co-design. For domestic AI infrastructure self-reliance, Zhenwu M890's cluster solution is a development worth continued tracking.