China's domestic chips run 2.8T Kimi K3 on day zero
Alibaba Cloud says its Zhenwu M890 supernode is the first Chinese system to run Moonshot's 2.8-trillion-parameter model; Huawei Ascend also claims day-0 support.
A supernode built for trillion-parameter inference
Alibaba Cloud said on Tuesday that its Lingjun Zhenwu M890 supernode instance has completed day-zero adaptation for Kimi K3, the 2.8-trillion-parameter open-weight model Moonshot AI released on Monday. The company describes it as the first system in China to successfully run a model of nearly three trillion parameters end to end.
The instance is built around 64 Zhenwu M890 accelerators from Alibaba's in-house chip unit T-Head, linked by a proprietary ICN Switch interconnect that the company says delivers 800GB/s of all-to-all bandwidth and pools roughly 9TB of high-bandwidth memory in a single coherent domain. That memory footprint is the binding constraint for a model of K3's size: without pooling on this scale, serving the weights requires spreading them across many loosely coupled nodes, and cross-node traffic dominates latency.
Alibaba Cloud, T-Head and Moonshot engineers jointly tuned the software stack and operator kernels for the model. Alibaba reports a roughly 35 percent reduction in time-to-first-token and a 1.8x improvement in per-accelerator decode throughput versus the unoptimised baseline. Alibaba also said its Qianwen AI platform and the Bailian model service will offer Kimi K3 APIs — notable given that Moonshot is a direct competitor to Alibaba's own Qwen family.
Huawei made a parallel claim the same day. Its computing division said Ascend achieved zero-day compatibility with K3, reproducing training runs on Atlas 800 A3 and Atlas 900 A3 SuperPoD systems with custom operators and parallelism optimisations, and serving inference through the open-source vLLM Ascend and SGLang backends. Huawei added that its Ascend 950 supernode natively supports the mxFP4 quantised weights Moonshot shipped, alongside FP8 and MXFP8 formats.
Why it matters
The significance is less about benchmark numbers than about the closing of a loop. Until recently, China's largest open models were trained and served predominantly on Nvidia silicon, leaving the domestic ecosystem dependent on export-controlled hardware for its most visible achievements. Two domestic vendors matching a frontier-scale release on launch day — one of them a rival lab's cloud arm — suggests the software maturity gap, historically the harder problem than raw FLOPS, is narrowing. It also gives Chinese enterprises a procurement path for K3-class models that does not route through Washington's licensing regime.