AI & Compute

DeepSeek and Huawei Build Open-Source Software Stack for Ascend

DeepSeek is open-sourcing compute and communication libraries for Huawei's Ascend chips and co-developed a 128-chip Ascend 950 supernode, pitching TileLang as a simpler alternative to Nvidia's CUDA.

By
Tom Whitfield
Filed
Channel
AI & Compute
Read
3 min read

Chinese AI startup DeepSeek has partnered with Huawei Technologies to build programming tools optimized for Huawei's Ascend AI chips, mounting a direct attack on the software moat that underpins Nvidia's dominance of the AI accelerator market.

DeepSeek announced the collaboration in a post on its official WeChat account, confirming that it is open-sourcing programming infrastructure for the Ascend platform, including compute and communication libraries. The two companies also jointly developed a "supernode" system built on 128 Ascend 950 chips, optimized for both computation and communication, with Huawei providing full support on the software work, DeepSeek said.

The announcement lands two weeks after Huawei unveiled its next generation of AI processors and supernode computing systems. At that launch, Huawei said it expects its AI systems to be widely used for model training next year — a timeline that now has a credible software partner attached to it.

TileLang as the CUDA counterweight

Central to the effort is TileLang, an open-source, high-level programming language for AI chips. DeepSeek's argument is structural: building an independent GPU software ecosystem first requires a universal language that is easy to program yet can still unlock the hardware's full performance. Without such a layer, developers remain locked into proprietary stacks.

The company pitched TileLang as offering "a simpler programming model" than CUDA, Nvidia's software platform for its chips. That claim targets the single most durable element of Nvidia's competitive position. Nvidia's hardware leads are periodically challenged — by Huawei's Ascend line, by domestic rivals, and by export-control erosion of its China market — but CUDA's two-decade accumulation of libraries, tools, and developer familiarity has kept switching costs high.

If TileLang and the open-sourced Ascend libraries genuinely lower the barrier for developers to extract performance from Huawei silicon, the calculus for Chinese AI labs and cloud providers changes. They would no longer need to route around US export controls to get usable training infrastructure; they could build on a domestic stack with domestic tooling.

Why the supernode detail matters

The 128-chip Ascend 950 supernode is more than a software demonstration. Supernode architectures tie large numbers of accelerators into a single coherent compute-and-communication domain, and getting the communication libraries right is typically the hardest engineering problem at that scale. Huawei's willingness to provide full support on the software side signals that this is a coordinated hardware-plus-software push rather than a lab experiment by a model developer.

The sequencing also matters. Huawei's processor and supernode launch came first, with the expectation of broad use for model training next year. DeepSeek's contribution — open-source libraries and a programming language — addresses the adoption gap that has historically separated Ascend's silicon capabilities from its real-world deployment in AI training workloads.

For Nvidia, the immediate commercial impact is limited; the company's shares slipped 0.72% on the day. The longer-term risk is cumulative: each layer of a working non-CUDA stack that Chinese developers adopt makes the domestic alternative stickier and narrows the window in which Nvidia can recapture China demand if export rules loosen.

The partnership is the latest confirmed step in a broader pattern of Chinese technology companies assembling a parallel AI compute ecosystem around Huawei's silicon. Whether TileLang's simpler programming model translates into developer adoption at scale will determine whether the Ascend platform moves from roadmap promises to widely deployed training infrastructure next year.

Original: s3.tradingview.com

Share this article:

More from Tom Whitfield

Tom Whitfield

Show full bio

Staff writer covering consumer brands and retail at Chip Dispatch.

113 articles

Related articles

« Previous articleNext article »