AI & Compute

Deepseek Open-Sources CUDA Alternative for Huawei Ascend Chips

Deepseek open-sourced TileLang, a CUDA alternative, plus compute and interconnect libraries for Huawei Ascend, validating the stack on a 128-chip Ascend 950 supernode.

By
Grace Kim
Filed
Channel
AI & Compute
Read
3 min read

Deepseek is releasing open-source programming tools for Huawei's Ascend AI chips, headlined by TileLang, a language the company positions as easier to work with than Nvidia's CUDA — and it has already proven the stack on a supernode of 128 Ascend 950 processors.

The Chinese AI developer announced the release on its official WeChat channel, and Reuters confirmed the details. The package includes libraries for computation and for moving data between chips, all published as open source. Huawei "fully supported" the work, according to Deepseek. Beyond the libraries, the two companies jointly optimized a supernode — a cluster of 128 Ascend 950 chips.

TileLang itself originated at Peking University, and Deepseek has used it in production for roughly a year. The company's argument is structural: anyone building an independent software ecosystem for AI chips first needs a universal language that is easy to program yet still extracts full hardware performance. Deepseek contends TileLang offers a simpler programming model than CUDA. The New York Times reports that Deepseek first tested the language on older Nvidia chips and that TileLang is now the company's main tool for its AGI research.

Why software decides the outcome

The partnership attacks the single largest obstacle facing China's AI hardware push. Domestic accelerators have advanced quickly, but they lack the software layer that squeezes performance out of the silicon.

Nvidia's dominance never rested on chip design alone. An estimated four million developers worldwide build on CUDA, and that ecosystem forms a moat that rivals such as AMD have failed to cross even when their hardware matched Nvidia's on paper. Huawei wants to close exactly that gap. Two weeks before Deepseek's announcement, the company unveiled new AI processors and supernode systems and said they would see wide use for model training next year — a roadmap commitment, not a shipped volume.

Capacity constraints shape the commercial picture too. Huawei admits it cannot meet domestic demand and plans to sell fewer chips abroad. Rotating chairman Eric Xu, referencing US export controls, said the company cannot accept a future that hinges on whether others are willing to sell chips to China.

How much of the moat is left

SemiAnalysis has been quantifying CUDA's remaining advantage, and its recent findings are blunt. After testing Jalapeño, OpenAI's inference chip, the analysts called the CUDA moat "potentially dead," because OpenAI ported new models to its own hardware remarkably fast. Jalapeño beat Nvidia's Blackwell on performance per watt in most tested scenarios, and SemiAnalysis notes OpenAI's own models — which run on Nvidia GPUs — helped design the chip.

The caveats matter. The analysts tested only scenarios that are relatively easy to optimize, around 8,000 input tokens and 1,000 output tokens. They have not yet run AgentX, the benchmark for multistep AI agent workloads. That is precisely where SemiAnalysis found Nvidia far ahead in an August analysis: with AMD's current software stack, Nvidia would remain cheaper per token even if AMD gave its hardware away for free. The analysts locate Nvidia's durable edge not in the silicon but in the software that binds many chips into a single system.

Huawei's chips did not appear in the AgentX comparison. In an earlier analysis of DeepSeek V4, however, SemiAnalysis observed that Huawei's CANN software stack was the only one besides CUDA to support the model on day one — a signal that the gap, while real, is narrowing at the edges.

For now, the Deepseek-Huawei release is a tooling announcement, not a displacement of CUDA in any measurable workload. But with TileLang open-sourced, a 128-chip Ascend supernode already optimized, and Huawei's new processors slated for broad training deployments next year, China's AI industry has assembled the first end-to-end alternative to the Nvidia-CUDA stack — and the AgentX benchmark results, when they arrive, will show how close it actually runs.

Original: mp.weixin.qq.com

Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering industry trends and analytics at Chip Dispatch.

97 articles

Related articles

« Previous articleNext article »