DeepSeek and Huawei open-source Ascend AI tools to rival Nvidia CUDA
New open-source compute, communication and kernel tooling makes Huawei’s Ascend chips more practical for large-scale AI training.
DeepSeek, in collaboration with Huawei, has open‑sourced a stack of programming tools for Huawei’s Ascend AI accelerators to reduce dependence on Nvidia’s CUDA ecosystem. The release includes DeepGEMM-Ascend for matrix and related tensor operations in BF16, FP8, and FP4, and DeepEP-Ascend for distributed communication, both built and validated on Ascend 950 hardware. TileLang, DeepSeek’s high‑level kernel language, now has native Ascend 950 support with codegen, auto‑scheduling, and synchronization, while still targeting Nvidia GPUs and other platforms. The libraries sit on top of Huawei’s CANN stack and were co‑designed and tuned for a 128‑Ascend‑950 supernode used to optimize computation and data movement for large‑scale training workloads. This work follows Huawei’s recent next‑gen Ascend and supernode announcements and extends the companies’ existing partnership around DeepSeek V4 and V4‑Flash models on Ascend chips.
Why it matters
By filling in core compute, communication and kernel language support on Ascend, DeepSeek and Huawei lower one of the main barriers to using non-Nvidia hardware for serious AI workloads. Teams that already use TileLang or DeepSeek’s libraries can now target Ascend 950 with familiar interfaces and features tuned on a large supernode, making it easier to treat Ascend as a first-class option alongside Nvidia GPUs.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- Tom's Hardware