
DeepSeek Open-Sources Ascend Software Stack, Targets CUDA Moat
DeepSeek open-sources TileLang, DeepGEMM and DeepEP ports for Huawei Ascend, with every V4 training operator now running on Chinese silicon as joint 128-card supernode work advances.
- By
- Tom Whitfield
- Filed
- Channel
- AI & Compute
- Read
- 4 min read
DeepSeek has open-sourced a full set of foundational training infrastructure for Huawei's Ascend platform, including a TileLang compiler toolchain that already covers every operator the company uses to train its V4 model family. The release, announced on DeepSeek's official Chinese WeChat account, pairs each new Ascend component with a direct counterpart in the company's previously released NVIDIA toolchain — a deliberate one-to-one mapping that signals an effort to make Chinese AI silicon a drop-in alternative rather than a one-off port.
The open-sourced projects cover the full low-level software stack needed for frontier-model training. TileLang, hosted at tile-ai/tilelang on GitHub, is a high-level language and compiler toolchain. DeepGEMM-Ascend accelerates general matrix operations. DeepEP-Ascend handles large-scale cross-device communication. TileKernels supplies vector computation and memory-access operators for data processing, FlashMLA provides sparse attention operators for long-context efficiency, and DeepSelect covers data selection.
According to DeepSeek, compute and communication performance of these components "is already approaching the limits of the underlying hardware" across a number of key test cases. The company did not publish benchmark figures comparing Ascend performance against NVIDIA GPUs.
The TileLang layer
TileLang is the centerpiece. DeepSeek first validated the language on NVIDIA's mature platform; it now implements the majority of operators used in training the DeepSeek V4 family of models and has become, in the company's words, "a core tool for our exploration of new AGI paradigms and development of high-performance operators."
The Ascend version released today wraps Huawei's low-level Ascend C instructions and exposes a high-level programming interface without sacrificing hardware performance. DeepSeek states that every TileLang operator currently used in its training workloads now has a corresponding high-performance implementation on Ascend.
The strategic significance lies in the abstraction. TileLang claims a simpler programming model than CUDA — one that "significantly improves development efficiency and simplifies code logic" — while still exploiting the characteristics of the underlying chips, whether NVIDIA or Ascend. If the same operator code runs efficiently on both platforms, the engineering cost of migrating AI workloads away from NVIDIA hardware drops substantially. That addresses the barrier that has constrained Huawei's Ascend line more than raw compute: the absence of a mature software ecosystem comparable to CUDA and its surrounding tooling.
Huawei's involvement runs deep
DeepSeek explicitly credits Huawei's team with "strong and wholehearted support" throughout the development. The two companies are jointly building a 128-card supernode solution based on the Ascend 950 and are conducting joint deep optimization of both computation and inter-chip communication. That collaboration goes beyond porting libraries; it extends into cluster-level systems design, where inter-chip communication efficiency — DeepEP's domain — determines how thousands of accelerators scale during training runs.
The partnership reframes the competitive picture. This is no longer a case of a model developer tolerating domestic chips as a fallback. DeepSeek and Huawei are jointly attacking the systems problem — silicon, interconnect and software together — that determines whether frontier models can be trained on Chinese hardware at competitive efficiency.
What it changes
For Huawei, the release gives Ascend a credible, openly available software layer maintained by one of the most closely watched AI labs in China. For other domestic accelerator vendors, DeepSeek says it hopes the open-sourced TileLang for Ascend will serve as "a useful reference for building highly usable software ecosystems around a broader range of AI chips" — an implicit invitation to standardize on the TileLang abstraction across non-NVIDIA hardware.
For NVIDIA, the development chips at the ecosystem lock-in that has underpinned its dominance in large-model training even where its GPUs are restricted. US export controls have already pushed Chinese labs toward domestic accelerators; DeepSeek's toolchain reduces the remaining friction — developer productivity and code portability — that kept CUDA entrenched.
DeepSeek says it will continue working with the community to build an open software ecosystem. With the Ascend 950-based 128-card supernode still under joint development and performance currently characterized only as "approaching hardware limits" without published cross-platform numbers, the next measurable signal will be whether a future DeepSeek frontier model trains end-to-end on Ascend at scale.
Original: substackcdn.com
More from Tom Whitfield
Show full bio
Staff writer covering consumer brands and retail at Chip Dispatch.
113 articles
Related articles
deepseek-open-sources-six-tools-to-build-software-stack-for-huawei-ascend-7ac77e33
DeepSeek Open-Sources Six Tools to Build Software Stack for Huawei Ascend
deepseek-and-huawei-build-open-source-software-stack-for-ascend-7fc48f63
DeepSeek and Huawei Build Open-Source Software Stack for Ascend
deepseek-open-sources-cuda-alternative-for-huawei-ascend-chips-37ccc5a0
Deepseek Open-Sources CUDA Alternative for Huawei Ascend Chips
deepseek-open-sources-huawei-ascend-tools-targeting-nvidia-s-cuda-moat-c777948b
DeepSeek Open-Sources Huawei Ascend Tools, Targeting Nvidia's CUDA Moat
