Compute#AI chips#Model inference

DeepSeek open-sources its Ascend stack: every V4-series operator now has a Huawei-chip port

DeepSeek published six Ascend component repos on September 30; it says every V4-series NVIDIA operator now has an Ascend counterpart.

Macro close-up of an oxidized silicon wafer

DeepSeek published six Ascend component repositories under its GitHub organization on September 30, completing the operator layer that runs its V4-series models on Huawei Ascend chips; Huawei posted authorized benchmark data the same day. It is the first time DeepSeek has released its chip-side software stack at this scale, after its model weights.

Facts

  • Six repos: DeepGEMM-Ascend (matrix compute), DeepEP-Ascend (cross-device communication; claimed Dispatch 375 GB/s, Combine 347 GB/s), TileKernels, FlashMLA (sparse attention) and DeepSelect, all under github.com/deepseek-ai.
  • TileLang: the high-level operator language (tile-ai/tilelang). DeepSeek says most V4-series NVIDIA operators are written in it, and every one of them has an Ascend counterpart — the basis of the “fully ported” claim.
  • CANN recipes: Huawei maintains companion docs on gitcode (cann-recipes-infer / cann-recipes-train / torchtitan-npu), covering inference communication, quantized training, LoRA and Agentic RL deployment.
  • Reported benchmarks: Huawei cites DeepSeek-V4.1-Flash offline inference (EP32, 128K context) at 2,469 tokens/s per card at 5 ms TPOT and 5,102 tokens/s at 10 ms, over a 128-card supernode with 3.2 Tbps scale-up.
  • Sourcing note: the QbitAI piece is a Huawei-authorized repost; treat repo documentation as the primary source. No independent third-party re-test exists yet.

Editorial take

Model releases are common; open-sourcing the operator layer is rarer. Training frameworks can be swapped, but the operator-chip coupling is the hardest thing to move, and publishing the full V4 port amounts to a reproducible endorsement of Ascend — a data point for anyone weighing China’s approval of Nvidia chip purchases. Wait for third-party re-tests on real workloads, and read it alongside the DSec training infrastructure DeepSeek disclosed the same week.