DeepSeek releases AI chip programming software developed with Huawei

Photo: Jakub Pabis / Pexels

DeepSeek releases AI chip programming software developed with Huawei

The open-source toolkit gives Chinese developers a homegrown alternative to Nvidia's CUDA ecosystem, marking a milestone in Beijing's push for tech self-sufficiency.

DeepSeek just handed Chinese AI developers something they’ve been waiting for: a full software stack purpose-built for Huawei’s chips, no Nvidia required.

The Chinese AI startup announced on September 30 via its WeChat account that it has open-sourced a comprehensive suite of programming tools designed specifically for Huawei Technologies’ Ascend AI accelerators. The toolkit includes TileLang, a high-level programming language that functions as a domestic counterpart to Nvidia’s CUDA, the software framework that has long served as the lingua franca of AI chip programming.

What’s in the toolkit

The release goes well beyond a single programming language. DeepSeek packaged together a collection of compute and communication libraries, including DeepGEMM, FlashMLA, TileKernel, DeepSelect, and DeepEP. Each addresses a different layer of the AI development process, from matrix multiplication to memory management to inter-chip communication.

The software is optimized for Huawei’s Ascend 950 chips and supports a “supernode” configuration linking 128 of those accelerators together.

Advertisement

The tools are available for free download, a deliberate choice to lower the barrier for developers who might otherwise default to Nvidia’s well-established ecosystem simply because switching costs felt too high.

Huawei reportedly provided extensive support throughout the development process.

The CUDA problem, and why this matters

To understand the significance of this release, you need to understand CUDA’s stranglehold on AI development. Nvidia’s programming framework isn’t just popular. It’s effectively a requirement. The vast majority of AI models, training pipelines, and inference systems worldwide are built on CUDA. Switching away from it means rewriting code, retraining teams, and rearchitecting workflows.

That dependency became a strategic vulnerability for China when the US imposed export controls restricting access to Nvidia’s most advanced chips. Beijing could design its own silicon through companies like Huawei, but without a mature software ecosystem to program those chips, the hardware risked becoming an expensive paperweight.

DeepSeek’s toolkit directly addresses that gap. TileLang gives developers a high-level abstraction layer similar to what CUDA provides, meaning programmers familiar with Nvidia’s workflow can transition to Huawei hardware without learning an entirely alien system. The accompanying libraries handle the performance-critical operations that AI workloads demand.

Building the infrastructure to match

Software alone doesn’t create an ecosystem. DeepSeek appears to understand this, with plans to construct a large data center in Inner Mongolia designed to host at least 160,000 Huawei Ascend accelerators.

The release also builds on groundwork DeepSeek laid earlier this year. In April 2026, the company adapted its V4 AI model to run on Huawei hardware, effectively proving that cutting-edge models could perform on non-Nvidia silicon. The new toolkit generalizes that effort, giving any developer the means to do the same with their own models.

Benchmarking tools included in the release assist developers with kernel development, providing standardized ways to measure performance and optimize code for the Ascend architecture.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
DeepSeek releases AI chip programming software developed with Huawei
DeepSeek releases AI chip programming software developed with Huawei

The open-source toolkit gives Chinese developers a homegrown alternative to Nvidia's CUDA ecosystem, marking a milestone in Beijing's push for tech self-sufficiency.

Photo: Jakub Pabis / Pexels

DeepSeek just handed Chinese AI developers something they’ve been waiting for: a full software stack purpose-built for Huawei’s chips, no Nvidia required.

The Chinese AI startup announced on September 30 via its WeChat account that it has open-sourced a comprehensive suite of programming tools designed specifically for Huawei Technologies’ Ascend AI accelerators. The toolkit includes TileLang, a high-level programming language that functions as a domestic counterpart to Nvidia’s CUDA, the software framework that has long served as the lingua franca of AI chip programming.

What’s in the toolkit

The release goes well beyond a single programming language. DeepSeek packaged together a collection of compute and communication libraries, including DeepGEMM, FlashMLA, TileKernel, DeepSelect, and DeepEP. Each addresses a different layer of the AI development process, from matrix multiplication to memory management to inter-chip communication.

The software is optimized for Huawei’s Ascend 950 chips and supports a “supernode” configuration linking 128 of those accelerators together.

Advertisement

The tools are available for free download, a deliberate choice to lower the barrier for developers who might otherwise default to Nvidia’s well-established ecosystem simply because switching costs felt too high.

Huawei reportedly provided extensive support throughout the development process.

The CUDA problem, and why this matters

To understand the significance of this release, you need to understand CUDA’s stranglehold on AI development. Nvidia’s programming framework isn’t just popular. It’s effectively a requirement. The vast majority of AI models, training pipelines, and inference systems worldwide are built on CUDA. Switching away from it means rewriting code, retraining teams, and rearchitecting workflows.

That dependency became a strategic vulnerability for China when the US imposed export controls restricting access to Nvidia’s most advanced chips. Beijing could design its own silicon through companies like Huawei, but without a mature software ecosystem to program those chips, the hardware risked becoming an expensive paperweight.

DeepSeek’s toolkit directly addresses that gap. TileLang gives developers a high-level abstraction layer similar to what CUDA provides, meaning programmers familiar with Nvidia’s workflow can transition to Huawei hardware without learning an entirely alien system. The accompanying libraries handle the performance-critical operations that AI workloads demand.

Building the infrastructure to match

Software alone doesn’t create an ecosystem. DeepSeek appears to understand this, with plans to construct a large data center in Inner Mongolia designed to host at least 160,000 Huawei Ascend accelerators.

The release also builds on groundwork DeepSeek laid earlier this year. In April 2026, the company adapted its V4 AI model to run on Huawei hardware, effectively proving that cutting-edge models could perform on non-Nvidia silicon. The new toolkit generalizes that effort, giving any developer the means to do the same with their own models.

Benchmarking tools included in the release assist developers with kernel development, providing standardized ways to measure performance and optimize code for the Ascend architecture.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.