Photo: Jakub Pabis / Pexels
DeepSeek releases AI chip programming software developed with Huawei
The open-source toolkit gives Chinese developers a homegrown alternative to Nvidia's CUDA ecosystem, marking a milestone in Beijing's push for tech self-sufficiency.
DeepSeek just handed Chinese AI developers something they’ve been waiting for: a full software stack purpose-built for Huawei’s chips, no Nvidia required.
The Chinese AI startup announced on September 30 via its WeChat account that it has open-sourced a comprehensive suite of programming tools designed specifically for Huawei Technologies’ Ascend AI accelerators. The toolkit includes TileLang, a high-level programming language that functions as a domestic counterpart to Nvidia’s CUDA, the software framework that has long served as the lingua franca of AI chip programming.
What’s in the toolkit
The release goes well beyond a single programming language. DeepSeek packaged together a collection of compute and communication libraries, including DeepGEMM, FlashMLA, TileKernel, DeepSelect, and DeepEP. Each addresses a different layer of the AI development process, from matrix multiplication to memory management to inter-chip communication.
The software is optimized for Huawei’s Ascend 950 chips and supports a “supernode” configuration linking 128 of those accelerators together.
The tools are available for free download, a deliberate choice to lower the barrier for developers who might otherwise default to Nvidia’s well-established ecosystem simply because switching costs felt too high.
Huawei reportedly provided extensive support throughout the development process.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
The CUDA problem, and why this matters
To understand the significance of this release, you need to understand CUDA’s stranglehold on AI development. Nvidia’s programming framework isn’t just popular. It’s effectively a requirement. The vast majority of AI models, training pipelines, and inference systems worldwide are built on CUDA. Switching away from it means rewriting code, retraining teams, and rearchitecting workflows.
That dependency became a strategic vulnerability for China when the US imposed export controls restricting access to Nvidia’s most advanced chips. Beijing could design its own silicon through companies like Huawei, but without a mature software ecosystem to program those chips, the hardware risked becoming an expensive paperweight.
DeepSeek’s toolkit directly addresses that gap. TileLang gives developers a high-level abstraction layer similar to what CUDA provides, meaning programmers familiar with Nvidia’s workflow can transition to Huawei hardware without learning an entirely alien system. The accompanying libraries handle the performance-critical operations that AI workloads demand.
Building the infrastructure to match
Software alone doesn’t create an ecosystem. DeepSeek appears to understand this, with plans to construct a large data center in Inner Mongolia designed to host at least 160,000 Huawei Ascend accelerators.
The release also builds on groundwork DeepSeek laid earlier this year. In April 2026, the company adapted its V4 AI model to run on Huawei hardware, effectively proving that cutting-edge models could perform on non-Nvidia silicon. The new toolkit generalizes that effort, giving any developer the means to do the same with their own models.
Benchmarking tools included in the release assist developers with kernel development, providing standardized ways to measure performance and optimize code for the Ascend architecture.