DeepSeek Open-Sources Ascend Kernels for V4.1 Inference
DeepSeek has released optimized attention kernels for Huawei’s Ascend 950, showing how open software can improve China’s alternative AI stack.
Software becomes the strategic layer
DeepSeek has open-sourced inference components for Huawei’s Ascend platform, including sparse-attention prefill and decoding kernels designed for the Ascend 950 NPU. The release is part of the FlashMLA repository and targets the infrastructure needed to run DeepSeek V4.1 efficiently outside NVIDIA’s CUDA ecosystem.
DeepSeek reports that the Ascend kernels reach up to 410 teraflops during prefill, equivalent to about 95 percent of the hardware’s theoretical peak, and 360 teraflops during decoding, or roughly 83 percent. The project also includes fused operations and a reported 10 to 15 percent decoding improvement for one kernel path.
These figures are vendor-reported measurements, not independent system-level benchmarks. They also come with constraints. The new release drops support for NVIDIA Hopper and older DeepSeek model versions, changes the FP8 and FP4 key-value cache format, and requires a specific Ascend software environment including CANN and torch_npu.
Why it matters
The significance is larger than one repository. High-end AI hardware is useful only when its software stack exposes enough of the silicon’s performance. By publishing optimized kernels, DeepSeek is contributing engineering knowledge that can lower the cost of deploying its models on Huawei hardware and make Ascend a more credible target for researchers and infrastructure teams.
The release also signals a deeper alignment between a major Chinese model developer and a domestic accelerator ecosystem. It does not prove that Ascend can match NVIDIA across workloads, tooling, or global availability. But it weakens the assumption that hardware substitution alone determines China’s AI constraints: sustained software optimization can materially change the economics of an alternative platform.