Dell and NVIDIA target rising AI costs with local compute hardware

Enterprises face mounting bills as agentic AI consumes significantly more tokens than traditional chatbots. New desktop units aim to shift workloads off the cloud to curb expenses.
The cost of running artificial intelligence models is climbing rapidly, creating a financial headache for many organizations. As companies move from simple chatbots to more complex agentic systems that perform multi-step reasoning, the number of tokens required per task has surged. This shift is driving up cloud API bills, with projections suggesting that AI coding costs could soon exceed the average salary of a software developer within two years.
To address this, Dell Technologies has introduced the Pro Max series, specifically the GB10 and GB300 models. These devices are powered by NVIDIA Grace Blackwell superchips and are designed to run high-performance AI workloads directly on local hardware rather than in the cloud. By keeping computation on-premises, businesses aim to gain better control over their spending while maintaining the speed needed for advanced AI applications.
Agentic workflows drive token consumption
The primary driver behind these rising costs is the adoption of agentic AI. Unlike traditional chatbots that respond to a single prompt, agents break down tasks into multiple steps, each requiring interaction with a large language model. Research indicates that these agents can consume between four and fifteen times more tokens than standard interactions. This complexity means that even a simple request can result in a significant spike in usage fees.
Industry analysts note that global spending on AI inference is expected to surpass spending on training by 2026. This trend highlights a structural shift in how companies pay for AI. Instead of a one-time license fee, organizations now face variable costs based on usage. For IT leaders, this creates a difficult budgeting challenge, especially when the volume of work grows unpredictably with each new project or employee adoption.
Local compute offers cost control
The Dell Pro Max GB10 is designed to handle models with 30 to 200 billion parameters, allowing users to run up to eight agents simultaneously on their own desks. By processing these requests locally, the device eliminates the per-token fees associated with cloud services. This setup is particularly useful for AI developers and data scientists who need frequent, heavy usage without incurring massive infrastructure strain.
For larger teams, the GB300 model provides data center-class capabilities, supporting models up to one trillion parameters and up to 150 concurrent agents. This scalability allows enterprises to keep pace with growing workloads without the need to constantly renegotiate cloud contracts. The trade-off, however, is the initial capital expenditure required to purchase the hardware and the ongoing maintenance of local servers, which must be weighed against the long-term savings on token costs.
Balancing capital and operating spend
According to reporting from GN technics/hardware (en-US), this approach represents a strategic shift in how businesses manage their AI budgets. While cloud-based solutions offer flexibility and lower upfront costs, the variable nature of token pricing can lead to unpredictable monthly bills. Local hardware flips this model, converting variable operating expenses into fixed capital expenses.
This method is not without its challenges. Companies must invest in physical infrastructure, manage power and cooling requirements, and maintain the hardware over time. Additionally, not all workloads are suitable for local processing, and some specialized models may still require cloud access. Nevertheless, for organizations with high-volume, predictable AI needs, the ability to control token consumption offers a compelling alternative to the current cloud-dominated landscape.






