ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
DeepSeek Open-Sources Its Ascend Infrastructure: TileLang Rivals CUDA, Cutting Costs for Domestic Compute

DeepSeek Open-Sources Its Ascend Infrastructure: TileLang Rivals CUDA, Cutting Costs for Domestic Compute

AI information • Admin • • 6 views

On September 30, 2026, DeepSeek open-sourced infrastructure components for Ascend. DeepSeek announced today via its official WeChat account that it is officially open-sourcing six infrastructure components for Huawei's Ascend computing platform: TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, each mirroring the components it previously open-sourced for the NVIDIA platform. All official repositories are on GitHub (organizations tile-ai and deepseek-ai).

What makes this release worth watching is not the models themselves, but the "software stack." The weakness of domestic AI chips has never been raw compute, but the ecosystem: the toolchains for writing code and tuning performance are far less mature than NVIDIA CUDA. Foreign media following the story see it as a key step for China's AI industry in closing its software-ecosystem gap.

Six components in one release

  • TileLang: a Pythonic domain-specific language and compiler toolchain built on the TVM compiler infrastructure, balancing developer productivity with low-level optimization. The GitHub repository tile-ai/tilelang announced in today's README update that it now officially supports native code generation for the Ascend 950 NPU, automatic scheduling and synchronization, and SIMD/SIMT vector programming.
  • DeepGEMM: a matrix-acceleration library; the Ascend version is at deepseek-ai/DeepGEMM-Ascend.
  • DeepEP: a distributed communication library; the Ascend version is at deepseek-ai/DeepEP-Ascend.
  • TileKernels: a collection of general-purpose kernels; FlashMLA: sparse-attention operators; DeepSelect: data-selection tooling.

DeepSeek also confirmed that the vast majority of operators used to train the V4 series were implemented in TileLang — these tools have already been validated in cutting-edge large-model training, not just on paper.

Why it's called "a CUDA counterpart"

The mapping is direct: in NVIDIA's ecosystem, CUDA is the programming-language layer, cuBLAS handles matrix acceleration, and NCCL handles collective communication. TileLang corresponds to CUDA's language layer, DeepGEMM to high-performance matrix computation, and DeepEP to collective communication. Once developers become fluent in this toolchain, they are no longer locked to a single hardware platform — "lower cost, higher efficiency" made concrete means domestic compute becomes usable and affordable.

The Huawei collaboration: near the theoretical performance ceiling

This is not a simple "code port." According to DeepSeek, Huawei's team was deeply involved in the co-development: together they implemented the 128-card supernode solution for the Ascend 950 and deeply tuned the compute and communication paths, with multiple tests showing compute and communication performance approaching the hardware's theoretical limits. Supernodes are the key form factor for large-model training — the Ascend platform is now battle-ready for flagship training runs.

Who benefits, and where the barriers are

The direct beneficiaries are teams training and running inference for large models on domestic compute: they can optimize operators and accelerate training without reinventing the wheel. There are two barriers: the hardware targets Ascend NPUs, so it's hard to get started without an Ascend environment; and the tech stack demands familiarity with both TileLang's DSL and low-level kernel tuning. But the bigger significance is the industry signal — when a top-tier model team starts building its toolchain on domestic hardware, the ecosystem flywheel starts turning.

Recommended Tools

More