Read as article
Ascend Supernode Pushed to Hardware Limits as Huang Concedes China
By @sharedot · · 9 pages
- Programming
- Open Source
- Deepseek
- Huawei Ascend
- Cuda
DeepSeek and Huawei report their 128-chip Ascend 950 supernode has been tuned to near hardware limits, while Nvidia's CEO says China is conceded.
What Happened: The Toolkit's Next Chapter
Following the September 30 open-source release of TileLang and five companion modules for Huawei's Ascend chips, new details have emerged about the depth of the engineering collaboration behind it. Tech Times reports that DeepSeek and Huawei jointly optimized a 128-chip supernode architecture built from Ascend 950 accelerators, tuning both computational workloads and inter-chip communication performance, with some test cases on the system approaching hardware performance limits. Huawei provided full engineering support throughout development, and the release was announced via DeepSeek's official WeChat channel. The package was designed to mirror one-to-one the open-source tools DeepSeek had previously released for Nvidia hardware, so developers can migrate workloads by changing a configuration flag rather than rewriting kernel code.
The Six-Module Stack in Detail
At the center of the release is TileLang, a domain-specific language DeepSeek now describes as its core tool for AGI research. Tech Times details the five supporting modules: DeepGEMM-Ascend handles matrix multiplication kernels and, per news.lavx.hu, supports BF16, FP8, and FP4 operations using the same programming interfaces as DeepSeek's existing DeepGEMM library. DeepEP-Ascend manages large-scale communication across devices, including routing data to experts in mixture-of-experts models. TileKernels covers vector operations and memory access, FlashMLA is optimized for long-context processing, and DeepSelect handles data filtering. TileLang now officially supports the Ascend 950 as a backend alongside Nvidia CUDA, AMD ROCm, and Apple Metal.
Why It's Surprising: Nvidia's Own Concession
The most striking development inside this story is Nvidia's reaction. Tech Times reports that in a May 2026 interview, Nvidia CEO Jensen Huang acknowledged that US export restrictions had cost the company roughly "$50 billion" in China this year, saying directly: "We have largely conceded the China market to Huawei." Nvidia's most recent annual filing also stated that competitors have built "larger developer and customer ecosystems to challenge us worldwide." For a company whose CUDA moat has held for nearly two decades, that admission reframes the DeepSeek-Huawei toolkit from a regional workaround into a direct challenge to the software ecosystem itself.
The Evidence Behind the Performance Claims
The supernode claims are not vague benchmark theater. Businesskorea reports DeepSeek's own statement that the two companies jointly pushed the computing and communication performance of the supernode solution to the hardware's limits, with Huawei providing extensive technical support throughout. The release also builds on Huawei's CANN software platform, which provides the infrastructure layer for running AI workloads on Ascend chips, according to news.lavx.hu.
The Stakes: The Training Gap Remains Open
DeepSeek's infrastructure remains bifurcated: Tech Times reports the company uses Huawei hardware for inference at scale and Nvidia hardware for training, and that a prior attempt to train a model on Huawei silicon stalled due to persistent technical difficulties, according to the Financial Times. DeepSeek has ordered 160,000 Ascend 950DT chips for its gigawatt-scale data center in Ulanqab, Inner Mongolia. Huawei rotating chairman Eric Xu said at Huawei Connect 2026 that Huawei expects Ascend systems to be "widely used for model training" in 2027 — an unverified forward-looking claim. Meanwhile, Tech Times notes SemiAnalysis's August 2026 AgentX benchmark found the CUDA moat remains decisive for multi-step agent tasks, though Ascend was not included in the comparison.
What Comes Next: A Community Flywheel Attempt
Tech Times frames the release's real question as whether open-sourcing a multi-backend abstraction layer can start the same community flywheel that built CUDA's moat after 2007 — each developer who adopts TileLang, contributes kernels, and publishes benchmarks makes it more useful for the next. Businesskorea reports Huawei has decided to limit overseas sales while it struggles to meet even domestic demand, so Chinese companies are expected to prioritize a robust closed ecosystem of domestic chips plus dedicated software. Industry observers expect the software release to lower the barrier to AI development in China; whether it shrinks the pool of developers for whom CUDA is the only reasonable choice will play out over years, not weeks.
V4 Day-Zero Adaptation Set the Stage
The September 30 release is the codification of work proven earlier in 2026. Tech Times reports that when DeepSeek V4 launched on April 24, 2026, four domestic chip vendors — Huawei Ascend, Cambricon, Hygon, and Moore Threads — confirmed full compatibility on release day, a "day-zero adaptation" previously exclusive to Nvidia. BAAI's FlagOS reportedly enabled V4 to run across eight domestic chip families within 24 hours, and news.lavx.hu notes Huawei said Ascend 950 supernodes fully supported V4 and that its chips trained part of the lighter V4-Flash model. First prove the frontier model runs on Huawei chips, then open-source the software that made it possible.
Sources
- techtimes.com › DeepSeek and Huawei Open-Source Toolkit Lets Developers Ditch CUDA Without Rewriting Code
- news.lavx.hu › DeepSeek and Huawei open-source Ascend programming tools to challenge Nvidia's CUDA dominance
- businesskorea.co.kr › DeepSeek, Huawei Join Forces to Challenge Nvidia's CUDA Dominance