NewsTradingSentimentCalendarCommunityBriefing
Tech

NVIDIA's New AI Chips Promise Faster Inference

By Tech Desk · 2026-09-17 · 2 min read
A dense rack of black server units with glowing blue indicator lights and thick fiber optic cables connecting them
Illustration: Tradingbird

NVIDIA's latest Vera Rubin system shows significant speed gains in early benchmarks, but the technology is still in preview and carries high infrastructure costs.

NVIDIA has released initial performance data for its new Vera Rubin NVL72 system, claiming it is significantly faster than its predecessor. The results come from the latest MLPerf Inference benchmark, a standardized test used to measure how quickly AI systems can process data. According to the company, the new hardware delivers up to 3.7 times the throughput of the previous GB300 model on specific visual language tasks.

However, these figures are based on preview submissions rather than final, widely available products. The technology represents a major step in hardware density and speed, but it also demands a specialized, high-power infrastructure that few organizations can currently deploy. For most businesses, the practical impact of this speed boost remains theoretical until the software ecosystem matures and the hardware becomes broadly accessible.

Speed Gains in Early Testing

The primary advantage of the Vera Rubin system is its ability to generate more text tokens per second. In the benchmarks, which included complex models like DeepSeek-R1 and Qwen3-VL, the new hardware outperformed the older GB300 units. This speed is achieved through improved hardware components and optimized software that allows the system to process more data simultaneously without sacrificing accuracy.

The performance jump is particularly notable in scenarios where the AI must reason through multiple steps before providing an answer. In these agent-based tests, the new system showed a massive lead over the previous generation. This suggests that as AI applications become more complex and interactive, the raw speed of the underlying hardware will become a critical factor in user experience and operational efficiency.

High Power and Scaling Limits

The trade-off for this speed is the physical infrastructure required to support it. The NVL72 system consists of a dense rack of server units that consume significant amounts of electricity and require advanced cooling. While the company reports high scaling efficiency, meaning the system performs well as more units are added, this efficiency relies on a tightly integrated network of high-speed cables and switches.

This level of integration means the system is not a simple plug-and-play solution for average data centers. It requires a specific architectural setup to handle the heat and data flow. For organizations considering an upgrade, the cost of retrofitting facilities to support these high-density units is a major consideration that may offset the savings from faster processing.

Software Optimization Drives Performance

Hardware alone does not determine the final speed of an AI system; software plays an equal role. The reported performance gains include significant improvements from updated software libraries that manage how data moves through the chips. These optimizations allow the system to use memory more efficiently, reducing the bottleneck that often slows down large language models.

According to GN auto tech/hardware: computing hardware, the collaboration between hardware design and software development is central to these results. The system uses a technique called disaggregated serving, which separates different stages of the AI processing to maximize efficiency. This approach is complex to implement and requires deep expertise, highlighting that the benefit of this technology is not automatic but depends heavily on how well it is configured.

Based on reporting by HPCwire, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories