Huawei Accelerates Ascend AI Roadmap With 960 Chips and 4,096 NPU Atlas SuperPoD

Huawei is accelerating its AI hardware roadmap, bringing the Ascend 960 generation forward while outlining annual upgrades through Ascend 970 in 2028 and Ascend 980 in 2029. The company is also expanding its strategy beyond individual accelerators, focusing heavily on massive SuperPoD systems where memory capacity, optical networking and tightly connected NPUs become just as important as the performance of each chip.

During Huawei Connect 2026, Huawei confirmed that Ascend 960DT is scheduled for Q1 2027, arriving 3 quarters earlier than its previous roadmap, while Ascend 960PR is planned for Q3 2027, 1 quarter ahead of schedule. Ascend 970 will follow in 2028 and Ascend 980 in 2029 as Huawei moves toward what it describes as a 1 generation per year development cycle.

Roadmap information presented around the event lists Ascend 960DT with up to 2 PFLOPS of FP8 compute and 4 PFLOPS of FP4, paired with 288 GB of HBM delivering 9.6 TB/s of memory bandwidth. Ascend 960PR is more heavily focused on inference, increasing FP4 compute to as much as 8 PFLOPS while using 192 GB of HBM. The later Ascend 970 and 980 designs are expected to continue increasing compute, memory bandwidth and interconnect performance, although Huawei has emphasized that specifications farther into the roadmap remain subject to development.

The larger shift is happening at the system level. Huawei introduced the Atlas 960E SuperPoD, which combines as many as 4,096 Ascend 960 NPUs under unified memory addressing and delivers a claimed 8 EFLOPS of FP8 compute with up to 1 PB of HBM capacity. It also introduces Huawei's Hi ONE near packaged optical technology, providing 7.2 Tbit/s of transmission capacity per optical engine. Around 5,500 Hi ONE units replace what Huawei says would otherwise require approximately 48,000 800G optical modules, reducing system power consumption by more than 550 kW.

Huawei claims the resulting architecture can reach 99.8% system availability while reducing the communication overhead that becomes increasingly expensive as AI clusters grow. Multiple systems can then be connected through its UnifiedBus architecture, with Huawei describing configurations supporting up to 512,000 NPUs and as many as 1 million NPUs when using a multi rail topology.

That direction builds on the company's existing Atlas SuperPoD 950 strategy. Huawei is effectively trying to reduce the importance of direct one to one accelerator comparisons by treating thousands of NPUs, memory resources, storage and networking as one tightly coordinated computing platform. This matters because individual Ascend accelerators are still generally viewed as behind NVIDIA's highest end hardware, an issue also reflected in reported Ascend 950 comparisons with NVIDIA GB300.

Huawei's roadmap makes its strategy increasingly clear. Instead of depending entirely on building an accelerator that beats NVIDIA chip for chip, it is attacking the problem at infrastructure scale. More NPUs, larger shared memory pools, faster optical connections and reduced communication overhead could compensate for weaker individual silicon in workloads that scale efficiently across thousands of processors.

The challenge is proving that this works outside Huawei's own demonstrations. Peak FP8 and FP4 figures do not reveal software efficiency, cluster utilization, power requirements or sustained performance during real model training. CANN, PyTorch support and Huawei's broader software ecosystem will therefore matter almost as much as Ascend 960 itself. If Huawei can keep large clusters productive while shipping them in meaningful volume, Atlas could become a much more serious alternative for Chinese AI infrastructure even without winning a direct accelerator specification battle.

Would you consider Huawei's system scale approach a credible alternative to NVIDIA, or does individual accelerator performance and CUDA software support still matter more?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

ASRock Z990 Motherboards Surface in Shipping Records Ahead of Intel Nova Lake

Next
Next

China Develops DUV GAA Transistors, but Full Process Still Faces 5 nm Class Density Limits