Alibaba runs 2.4T-parameter Qwen3.8 on its own Zhenwu chips
Alibaba Cloud says its Zhenwu M890 supernode, built on in-house T-Head chips, now runs 2.4T-parameter Qwen3.8 in production, cutting inference costs over 40%.
Alibaba Cloud says its flagship model now runs in production on its own silicon. The company announced that Zhenwu M890, a supernode in its Lingjun computing line built around a unified training-and-inference chip from in-house designer T-Head (Pingtouge), is serving Qwen3.8 — Alibaba's 2.4-trillion-parameter mixture-of-experts flagship — for inference on the Bailian platform. Chinese outlets describe it as the first domestic supernode to serve a model above two trillion parameters.
The hardware
Each M890 supernode links 64 of the T-Head chips over an 800GB/s interconnect and pools 9TB of memory, with precision support spanning FP32 down to FP4. Alibaba claims the system delivers a 1.5x speedup on agentic inference workloads and cuts inference costs by more than 40 percent. The announcement drew broad coverage in Chinese tech and financial media but barely registered in Western feeds.
Vertical integration, Chinese edition
The move mirrors Google's TPU playbook: own the model, own the chip, own the cloud, and stop paying the Nvidia margin on inference. For Alibaba the strategic weight is heavier. US export controls already bar top-end Nvidia hardware from China, and Washington floated sanctions on Chinese AI models just last week. Running flagship-scale inference on home-grown silicon — at production scale, for paying cloud customers — is precisely the capability those policies were meant to slow.
Why it matters: inference, not training, is where the bulk of AI compute spending is heading, and Alibaba has now demonstrated that a 2-trillion-parameter-class model can be served commercially without American accelerators. If the claimed economics hold, expect other Chinese hyperscalers to accelerate their own supernode timelines — and expect the export-control debate in Washington to shift from whether China can get chips to whether it still needs them for the workloads that make money. The open question is training: serving Qwen3.8 on Zhenwu is one thing; training its successor there would be the real independence milestone.