Why AI Chips Are Switching to Shorter Memory Stacks

The industry is moving away from taller memory stacks to cut costs and solve supply shortages.
The semiconductor industry is undergoing a significant shift in how it builds high-performance AI accelerators. For the past few years, the standard approach has been to pack as much memory as possible into each chip. This involved using taller stacks of High Bandwidth Memory, or HBM, to increase data capacity and speed. However, this strategy is reaching a breaking point. Major chip designers are now opting for shorter memory stacks, a move that contradicts previous expectations of ever-increasing density.
This change is driven by a combination of supply chain constraints and economic realities. As reported by GN technics/hardware (en-US), the scarcity of HBM wafers has become a critical bottleneck. By using shorter stacks, manufacturers can produce more memory units from the same limited supply of raw materials. This strategic pivot aims to balance the need for high performance with the practical limits of production capacity and cost.
Supply Shortages Force Strategic Shift
The primary driver for this change is the severe shortage of HBM wafers. These wafers are the raw material used to create the memory chips. For a long time, the industry assumed that demand would continue to outpace supply, leading to a push for more complex and taller stacks. But the current production limits mean there simply are not enough wafers to support the previous design goals.
Moving to shorter stacks allows companies to maximize the number of usable memory units they can produce. For example, a major chipmaker recently adjusted its next-generation accelerator to use eight-layer stacks instead of the previously expected twelve or more. This decision ensures that more accelerators can be built with the available resources, even if each individual unit holds slightly less total memory. It is a pragmatic response to a physical constraint in the supply chain.
Cost Efficiency Drives New Standards
Beyond availability, there is a strong financial argument for shorter stacks. In many AI applications, particularly those focused on inference, the speed of data access matters more than the total amount of storage. Shorter stacks offer a better cost-to-bandwidth ratio, meaning they deliver high performance at a lower price per unit of data processed.
This approach helps reduce the overall cost of running AI models. By focusing on efficiency rather than raw capacity, companies can lower the cost per token generated. This is a significant trade-off, as it prioritizes speed and affordability over maximum storage. For many data centers, this balance is more valuable than having the largest possible memory footprint, especially when power and cooling costs are already high.
Balancing Performance and Practical Limits
The shift to shorter stacks represents a maturation of the AI hardware market. It moves the industry away from a simple race for maximum specifications and toward a more nuanced understanding of what is actually needed for different workloads. This change benefits both chip suppliers and their customers by creating a more sustainable and cost-effective ecosystem.
While this means some potential capacity is left unused, the overall system becomes more reliable and affordable. The catch is that developers must optimize their software to work within these new constraints. However, the result is a more stable supply of AI hardware and a reduction in the extreme price volatility that has plagued the memory market recently.






