Press Esc to close

DeepSeek teamed up with Huawei to attack the part of Nvidia that's hardest to copy

DeepSeek and Huawei: open-source chip tooling for Ascend, challenging Nvidia's CUDA software ecosystem

Everyone keeps score in the chip war by counting silicon. Who taped out what, how many flops, whose interconnect is fatter. That scoreboard misses the part that actually decides winners.

The part is software. Nvidia's real moat was never a single GPU. It was CUDA, the programming platform that a generation of researchers learned first and never had a reason to leave.

If you want to break that grip, you do not start with a faster chip. You start with tools developers might actually enjoy using. That is what makes this week's news interesting.

DeepSeek said on Wednesday that it has partnered with Huawei to build programming tools for Huawei's Ascend chips, and that it is open-sourcing the lot: compute libraries, communication libraries, the works. The announcement came through the company's official WeChat account, as Reuters reported.

Read the fine print and it is more than a press release friendship. Huawei gave what DeepSeek describes as full support on the engineering side, and the two jointly tuned a "supernode" of 128 Ascend 950 chips, squeezing both the math and the chatter between chips.

The chatter matters. Training a giant model is less about raw speed than about keeping hundreds of chips from standing around waiting on each other. Whoever fixes the waiting wins.

Timing is doing some work here too. The announcement landed two weeks after Huawei showed off its next generation of AI processors and said it expects them to carry real training workloads next year, as Reuters covered at the time.

None of that is the most interesting line in the announcement. The most interesting line is a programming language called TileLang.

TileLang is open source and pitched as a simpler way to write fast code for AI chips. DeepSeek's claim is blunt: easier to program than CUDA, without giving up the performance ceiling.

Strip away the diplomacy and the goal is stated plainly. China wants a GPU software stack it controls, and the first step is a high-level language general enough to catch on and low-level enough to be fast.

Sound familiar? It should. That is exactly how CUDA won: meet programmers where they are, then make the fast path the easy path. TileLang is an attempt to run the same play for Ascend.

DeepSeek is a plausible team to try it. For two years the lab has shipped models that punch far above their compute budget, and it has a habit of publishing its work instead of hoarding it. We covered the latest of those releases, DeepSeek V4.1 Pro, just days ago.

The two companies are not strangers either. Back in April, Huawei said its Ascend supernode would run DeepSeek's V4 models from day one, Pro and Flash versions included, with low-latency inference to show for it.

So this partnership has receipts. What changed is the layer: from running models on Ascend to making Ascend pleasant to program.

Each side gets something concrete. DeepSeek gets a hardware home that does not depend on supply chains it cannot touch. Huawei gets the model shop whose name carries weight with developers, plus tooling that could make Ascend the default choice for the next team that cannot buy Nvidia.

Open-sourcing the tools is the shrewd part. Free, good tooling is the cheapest customer acquisition a chip company can buy.

The hard question is whether developers will move. CUDA's advantage was never just technical. It is millions of lines of existing code, tutorials, forum answers, and muscle memory. Every new project inherits all of it for free.

TileLang is asking programmers to start over on hardware that is still proving itself. That is a big ask, however elegant the language.

Huawei's answer is a date: real training workloads next year. If Ascend clusters start showing up in training runs that matter, the tooling story gets much easier to believe.

Nvidia, for its part, is not standing still on the software front either. The company keeps extending its own platform, including a recent push around agent safety tooling. The moat gets deeper while challengers dig around it.

Still, there is a reason to take this seriously. DeepSeek could have kept these libraries private and enjoyed a quiet edge on Ascend. Instead it published them, which reads like an invitation: if you are writing for anything that is not Nvidia, start here.

Invitations are cheap. Adoption is everything. But as of this week, the code is public, and public code is very hard to unring.

Comments