Huawei Redesigns Cloud Infrastructure for AI Agents

Huawei argues that traditional cloud setups are too rigid for the next generation of autonomous software, proposing a new architecture that treats AI agents as core workers rather than add-ons.
At its recent summit in Shanghai, Huawei presented a significant shift in how enterprise cloud infrastructure is built. The company argues that the rise of agentic AI—software that acts independently to complete tasks—requires a fundamental redesign of the underlying technology. Instead of treating AI as a separate application layer, Huawei is integrating it directly into the core of its cloud stack, aiming to provide a more deterministic and efficient foundation for businesses.
The core of this strategy is a move away from traditional IT resource management, which focuses on separate silos of compute, storage, and networking. In the new model, the focus shifts to managing AI-specific resources like models and data tokens as a unified system. This approach is designed to prevent the fragmentation that often occurs when different departments manage their own AI tools, ensuring that resources are available on demand for any task an agent needs to perform.
Shifting Focus to Autonomous Tasks
Antonony Gu, President of Huawei Hybrid Cloud, outlined three key changes driving this transformation. First, the primary actors in business execution are changing from humans operating applications to AI agents collaborating with humans and other agents. Second, resource provisioning is becoming task-oriented, where compute and data are packaged for immediate use by these agents. Third, the management focus is expanding to include AI compute and model governance, which traditional systems do not handle well.
According to GN auto tech/cloud: cloud infrastructure, this shift means that cloud platforms must evolve from static resource providers into dynamic hubs that support continuous evolution. The new architecture includes an agile cloud infrastructure that unifies general-purpose and AI compute, allowing systems to scale and adapt as agent workloads change. This is crucial for enterprises looking to deploy AI at scale without running into bottlenecks or security gaps.
Improving Hardware Efficiency and Data Flow
A major practical benefit of this new stack is improved hardware efficiency. Huawei claims that by using technologies like NPU pooling and memory snapshots, compute utilization can rise significantly from roughly 30% to 70%. This means businesses can get more out of their existing hardware investments. Additionally, the platform introduces a decoupled storage and compute architecture, which allows data to be managed and accessed more flexibly, turning raw data into a productive asset for AI agents.
The system also emphasizes end-to-end security and observability. By integrating security into the entire lifecycle of AI agents, from development to deployment, the platform aims to mitigate risks associated with autonomous software. This includes providing a framework for governing how agents interact with data and models, ensuring that the increased productivity does not come at the cost of stability or safety.
Trade-offs in Complex Integration
While the potential for increased efficiency is clear, adopting such a comprehensive architecture requires a significant commitment. The move to an agent-native platform involves integrating multiple complex capabilities, including specialized foundation models for vision, scientific computing, and prediction. This depth of integration can make the initial setup and migration process more challenging for companies that have invested heavily in legacy systems.
Furthermore, the reliance on a unified ecosystem for data, compute, and models creates a tighter coupling between different parts of the IT stack. For enterprises, this offers the promise of seamless AI adoption but also introduces a dependency on a single provider’s roadmap and standards. The trade-off is that while the system is designed to be deterministic and efficient, it demands a level of architectural alignment that may not be feasible for all organizations immediately.






