A research team including Huawei Technologies has successfully used the company’s own Ascend 910C chips to complete full-parameter post-training for DeepSeek-V4-Pro, a 1.6 trillion parameter model. The cluster ran at least 1,000 Huawei chips and completed over 1,500 training iterations without a single interruption — the first documented case of Chinese domestic silicon handling complex model training, not just inference.
🔍 THE BOTTOM LINE
US export controls were designed to keep China dependent on Western chips for frontier AI. That assumption is now cracking. If Huawei’s Ascend 910C can reliably post-train a 1.6 trillion parameter model, the gap between Chinese and US silicon is narrower than the export control framework assumes — and it is closing faster than anyone in Washington predicted.
What Changed
The distinction matters: inference is running a model. Training is building one. Chinese chips have been able to handle inference for years — you can run a model on Huawei silicon if someone else trained it on Nvidia GPUs. But post-training, which refines a model’s weights after initial pre-training, requires sustained, high-throughput compute and sophisticated interconnect. That is the capability that has been missing.
The research team — Huawei, the Shenzhen Loop Area Institute, the Shenzhen campus of Harbin Institute of Technology, and the Shenzhen Research Institute of Big Data — ran full-parameter post-training, meaning the entire model architecture was updated, not a subset. The Shenzhen government announced the result on Friday. The model completed more than 1,500 iterations with zero errors, and the process improved the model’s mathematical capabilities.
As the Shenzhen announcement explained: previously domestic compute was “much like building a one-way road for the model: input a question, output an answer.” Post-training added “complex flyovers and loops to that one-way road, instantly multiplying the computational and communication demands by several times.”
Context
DeepSeek-V3 was trained on a cluster of 2,048 Nvidia H800 processors — chips that are now restricted under US export controls. When DeepSeek-V4 launched in April as open-source, local chip firms including Huawei, Moore Threads, and Cambricon announced “day-zero” inference compatibility. But DeepSeek has remained tight-lipped about what hardware V4 was trained on from scratch.
This post-training run on Huawei silicon is not the same as full pre-training. Pre-training — teaching a model to speak from scratch — requires far more compute and months of sustained throughput. Post-training teaches an already-pre-trained model to follow instructions, obey safety rules, and handle specific tasks. It is still computationally intensive, especially at 1.6 trillion parameters, but it is a step short of the full pipeline.
The distinction matters less than the trajectory. In April, domestic chips could only do inference. In June, they can do full-parameter post-training. The path from here to pre-training is engineering, not theory.
Jensen Huang’s admission that Nvidia has conceded China’s chip market to Huawei now looks less like a concession and more like a forecast.
NZ Angle
New Zealand sits at the edge of a technology cold war with no skin in the chip game and all the skin in the compute game. NZ research institutions and startups depend on cloud compute from US providers (AWS, Azure, GCP) that run on Nvidia silicon. If the bifurcation of AI compute into US-aligned and China-aligned spheres accelerates, NZ’s options narrow. The country has no domestic chip industry and no path to one. What it does have is diplomatic positioning — the ability to buy from both sides, if both sides are willing to sell.
The Huawei milestone is also a signal for NZ’s sovereign AI conversation. If China can train frontier models on domestic chips, the assumption that AI sovereignty requires US hardware is no longer safe. It does not mean NZ should buy Huawei — it means the vendor landscape is diversifying, and policy should reflect that.
The Other Side
One successful post-training run is not parity. The 1,000-chip cluster that ran DeepSeek-V4-Pro’s post-training is a research setup, not a production pipeline. Pre-training from scratch requires substantially more compute, sustained for months. Huawei’s Ascend chips still lag Nvidia’s latest Blackwell architecture in raw performance, and the software ecosystem — CUDA’s dominance is not just about hardware, it is about every library, tool, and optimization built on top of it — remains a moat.
The Shenzhen government’s announcement is also a government announcement, not a peer-reviewed paper. The claims of 1,500 error-free iterations and improved mathematical capabilities are self-reported. They may be accurate, but the incentive to overstate is high.
Still: the difference between “China cannot train frontier models” and “China can post-train but not pre-train” is a difference of timeline, not direction. The export control framework was built on the assumption that the gap was generational. It may be a year or two.
The Bigger Picture
Baidu’s executive vice president Shen Dou said last month that a major version of Ernie 5.1 was trained on a cluster powered by Baidu’s Kunlunxin chips. Meituan invited users to test a trillion-parameter model that local reports say was trained entirely on domestic chips. The pattern is not one breakthrough — it is a wall of announcements, each one pushing the boundary further.
DeepSeek’s recent multimodal expansion showed the company is not standing still on capabilities. Add domestic training to the mix, and the premise of export controls — that they buy time for the US to stay ahead — starts to look like a bet that may not pay off.
❓ FAQ
What is the difference between pre-training and post-training?
Pre-training teaches a model to predict the next token from massive data — it is how a model learns language and basic reasoning. Post-training refines that model to follow instructions, answer questions, and behave safely. Both are compute-intensive, but pre-training requires far more sustained throughput.
Does this mean DeepSeek-V4-Pro is now fully Huawei-trained?
No. The post-training was done on Huawei chips. The pre-training — the bigger, longer phase — was likely done on a different hardware stack that DeepSeek has not disclosed. This is a step toward domestic training, not the full distance.
Are US export controls failing?
It depends on the metric. If the goal was to prevent China from ever catching up, they are failing. If the goal was to slow the catch-up and buy time, they are working — but the time bought appears to be measured in months, not years.
What should NZ policymakers take from this?
Diversification. If AI compute is splitting into US-aligned and China-aligned stacks, NZ needs to maintain relationships with both. Betting entirely on US cloud providers looks less safe when the alternative stack is proven, not theoretical.
🔍 THE BOTTOM LINE
Huawei training a 1.6 trillion parameter model on domestic chips is a milestone that will be dismissed as insufficient by some and overhyped by others. The honest read is that the gap between US and Chinese AI silicon has narrowed from “insurmountable” to “closing.” The export control era was built on the first assumption. The second one requires a different playbook.