Huawei Atlas 960 Targets 10 Trillion Parameter AI Models

Huawei's new system uses a unique interconnect to link thousands of chips, aiming to rival Nvidia in large-scale AI training.
Key points
- Huawei's Atlas 960 system supports AI models with up to 10 trillion parameters using 15,488 processors.
- UnifiedBus interconnect links thousands of chips to reduce data communication bottlenecks in AI training.
- New optical technology reduces power consumption by over 550 kilowatts compared to conventional modules.
Huawei is deploying its Atlas 960 SuperPoD system to support artificial intelligence models with up to 10 trillion parameters. This expansion marks a significant shift in strategy, moving from competing on single-chip speed to leveraging massive system scale and improved data flow.
As reported by TechRadar, the company is addressing the limitations of individual accelerators by connecting thousands of processors into a unified network. This approach allows the hardware to handle complex training workloads that would overwhelm smaller, isolated setups.
System scale drives performance gains
The Atlas 960 roadmap outlines configurations with up to 15,488 processors, a major increase from previous designs. These systems are specified to deliver 30 EFLOPS of FP8 computing power and include 4,460 TB of memory. Huawei expects these setups to process 15.9 million tokens per second during training.
The company has also accelerated its product timeline, with the Ascend 960DT and 960PR chips now scheduled for 2027. This one-generation-a-year cycle aims to keep pace with industry demands while providing a clear path for future hardware upgrades.
UnifiedBus reduces communication bottlenecks
A core component of this architecture is UnifiedBus, an interconnect designed to link processors, memory, and storage with lower communication overhead. In the Atlas 960E variant, this system connects up to 4,096 NPUs, allowing them to access shared memory across the entire platform.
This design is critical because large AI models require continuous data exchange between chips. By reducing the friction in this data flow, Huawei compensates for the fact that individual Ascend chips may not match the raw speed of Nvidia’s leading products.
Optical tech cuts power usage
Huawei integrates Hi-ONE optical technology to improve data transmission speeds and reduce energy consumption. The Atlas 960E uses approximately 5,500 Hi-ONE units instead of 48,000 conventional optical modules. This change is projected to save more than 550 kilowatts of power in the interconnection system.
The trade-off for this efficiency is increased system complexity. While the architecture supports massive configurations of up to 512,000 NPUs, it requires a highly coordinated environment. This makes the system powerful for large-scale training but less flexible for smaller, standalone applications.






