Huawei Shifts AI Focus from Raw Compute to Network Speed

As AI models grow larger, the bottleneck is no longer just processing power but how quickly data moves between chips. Huawei is prioritizing networking hardware to unlock new performance levels.
The industry has spent the last few years chasing raw computing power, assuming that more chips would automatically solve AI challenges. However, as clusters expand to include tens of thousands of units, a hidden problem has emerged: the data cannot move fast enough. When models reach trillions of parameters, the time spent waiting for data to travel between processors becomes a major drag on efficiency. In training scenarios, this latency can consume up to half of the total processing time. In inference, where users wait for a response, network congestion directly causes the lag people feel when an AI seems slow.
Huawei is addressing this by shifting its strategy from simply adding more compute to optimizing the networking infrastructure that connects these chips. Their new approach treats the network as a primary driver of performance rather than just a supporting utility. By focusing on high-speed network interface cards and storage units, the company aims to ensure that data flows smoothly across massive server arrays. This shift is significant because it changes how AI infrastructure is designed, moving away from isolated processing units toward a highly interconnected system where speed of transfer is as critical as the speed of calculation.
Network latency becomes the primary bottleneck
Traditional logic suggested that as long as the processors were powerful enough, the system would perform well. That assumption is no longer valid. In modern large-scale AI tasks, communication between nodes accounts for a massive portion of the workload. Research indicates that in distributed training, network communication time can represent 30 to 50 percent of the total duration. For inference tasks, which power real-time applications, collective communication can exceed 30 percent of the load. Even minor fluctuations in this data flow can be perceived by users as severe lag, creating a direct link between network stability and user experience.
New hardware targets high-speed data flow
To combat these delays, Huawei has introduced the SP560 series, a network interface card designed for massive scale. This hardware supports a bandwidth of 800 Gbps, which is a significant leap from previous generations. It is built to handle connections across up to 100,000 nodes, using a dual-bus architecture that maximizes data throughput. The goal is to create a digital highway where data can move quickly enough to keep the processors busy, rather than waiting for instructions or results. This hardware is specifically engineered to reduce the idle time that occurs when chips are waiting for data to arrive.
Storage replaces some computation needs
Beyond networking, Huawei is also leveraging storage to improve performance. As AI models handle longer contexts, they generate large amounts of temporary data known as KV Cache. Instead of recalculating this data repeatedly, which consumes computing power, the new system uses dedicated storage units to keep this data readily available. This approach, described as using storage in place of computation, reduces the load on the main processors. It allows the system to retrieve previous results instantly, speeding up response times without requiring additional processing cycles. This trade-off prioritizes memory access speed over raw calculation power, a shift that optimizes the overall efficiency of the AI infrastructure.






