Big Tech

DeepSeek and Huawei Take Aim at Nvidia’s CUDA Advantage With Open-Source Tools for Ascend Chips

Zaara Abbas

By: Zaara Abbas

4 min read

DeepSeek has open-sourced programming tools built with Huawei for its Ascend AI chips, including compute and communication libraries, and the two companies have tuned a 128-chip Ascend 950 system together. The work targets the software that keeps developers tied to Nvidia, which matters as much as the chips themselves in China’s push for self-sufficiency.

[For more news, click here] 

The last time DeepSeek rattled Nvidia investors was in January 2025 because of its low-cost R1 model that helped knock close to $600 billion off the chipmaker’s market value in a single trading day. The Chinese AI lab’s latest move was a post on its official WeChat account on Wednesday. While it may not seem like much, it goes after something Nvidia has guarded for nearly two decades: the software developers use to program its chips.

DeepSeek said it had worked with Huawei Technologies to build programming tools tuned for Huawei’s Ascend AI chips and was open-sourcing that infrastructure, including compute and communication libraries. According to Chinese media accounts of the post, the release mirrors libraries DeepSeek had already published for Nvidia hardware, such as DeepGEMM for matrix math and DeepEP for moving data between chips, so any team building on Ascend can now download the same kind of tools DeepSeek uses on Nvidia systems.

The two companies also jointly developed a “supernode” system built from 128 Ascend 950 chips, tuning it for both computation and communication, and DeepSeek said Huawei gave the project its full support. The announcement came two weeks after Huawei unveiled its next generation of AI processors and supernode systems and said it expected them to be widely used for model training next year.

Why Software Matters as Much as Chips in the Race with Nvidia

Nvidia’s grip on AI has rested as much on CUDA, the programming platform it introduced in 2006, as on its processors. Millions of developers have learned it, most AI frameworks are tuned for it, and moving to another company’s hardware usually means rewriting and retesting code that already works. That cost of leaving is a big reason Chinese companies kept buying Nvidia chips even as Washington tightened export rules and Beijing nudged them toward homegrown alternatives.

DeepSeek’s rests on TileLang, an open-source programming language for AI chips that it says speeds up development and makes code easier to follow.

“To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential,” DeepSeek said. “TileLang was created precisely to meet this need,” the company added, saying the language offers “a simpler programming model” than CUDA.

In principle, an engineer writes a piece of code once in TileLang and the compiler works out how to run it well on each type of chip, whether it comes from Nvidia or Huawei. If that holds up at scale, the cost of switching drops sharply. DeepSeek has good reason to want it to work, since most of the operators used to train its V4 models are already written in TileLang, according to reports on the release.

What the DeepSeek and Huawei Partnership Means for Nvidia

For policymakers in Washington, the partnership shows the limits of export controls that focus on hardware. The US has tightened and loosened rules on which Nvidia processors can be sold into China several times over the past three years, and Beijing has at times told its companies to steer clear of them altogether, while Huawei says demand at home for its chips already runs ahead of what it can produce, according to reports of its September launch.

Huawei’s individual chips still trail Nvidia’s best on raw performance, which is why the company has leaned so heavily on wiring hundreds or thousands of them together into supernodes that behave like one large machine. For now, however good it is, persuading engineers outside DeepSeek to learn a new language will also take years because CUDA’s advantage comes from habit and a huge library of existing code as much as from technical merit. Huawei says around 5,200 developers work in its software ecosystem each month and that about 40 major AI models already train on its platform, according to reports of its launch event, figures that still look small next to the millions who write code for Nvidia.

DeepSeek has its own commercial reasons to want this to work, having hired CITIC Securities to prepare a listing in Shanghai, as reported earlier this month. A software stack that runs well on Chinese chips would leave the company, and anyone investing in it, far less exposed to swings in US trade policy.

Huawei, for its part, has pulled the launch of its next chip, the Ascend 960, forward to the first quarter of 2027, three quarters ahead of its original schedule, according to coverage of the September event. By then, the code DeepSeek released on Wednesday will already be sitting on public repositories, free for anyone building on Ascend hardware to use.

Related Articles

DeepSeek Taps CITIC Securities to Prepare a Shanghai IPO as Chinese AI Firms Rush to List

Why AMD is Buying Fei-Fei Li’s World Labs for $8.2 Billion to Compete in Physical AI

Why SK Hynix is Talking to Intel About Chips Made in America

Share this article

Related Articles