Key points
- Memory had reached 41% of a Rubin Ultra rack's capital cost
- Halving HBM per GPU takes that to 28%, SemiAnalysis estimates
- The savings move to scale-up networking, 4% to 12%
- SemiAnalysis calls it a supply-crunch symptom, not weak demand
SemiAnalysis put a number on why Nvidia (NVDA) stripped memory out of its next flagship AI system. After this year's price increases for high-bandwidth memory and DRAM, memory had grown to about 41% of the total capital cost of a Rubin Ultra rack, the research firm said in a post on X at 9 p.m. ET on Tuesday, September 1, after the market closed. Cutting the memory per chip in half brings that share down to 28%.
The change itself isn't new. The Information reported on August 5 that Nvidia was testing versions of Rubin Ultra with less memory than the 1 terabyte per package Jensen Huang described at GTC in 2025, and SK Hynix lost about 15% in Seoul over the two sessions that followed. SemiAnalysis itself said on June 30 that Nvidia had cut Rubin Ultra from four compute dies to two. The configuration SemiAnalysis now treats as settled uses HBM4 stacked eight dies high, for 192 gigabytes per GPU, in place of the HBM4E 12-high stacks that would have carried 384 gigabytes. That's below the 288 gigabytes on the regular Rubin chip that comes first.
What's new is the cost math, which the firm said it presented at a recent conference keynote. "Why strip memory out of your flagship rack-scale system? Because after the latest HBM and DRAM price hikes, memory had quietly become ~40% of total capital cost of ownership," the post said. The table SemiAnalysis published compares the two designs on a cost-per-GPU-hour basis and assumes a 70% markup on HBM and a 60% markup on DRAM for the GPU provider.
| Cost layer | Rubin Ultra HBM4E 12-Hi (NVL72) | Rubin Ultra HBM4 8-Hi (NVL572) |
|---|---|---|
| HBM | $0.99 per GPU-hour (29%) | $0.43 (14%) |
| DRAM | $0.26 (7%) | $0.26 (9%) |
| NAND | $0.16 (5%) | $0.16 (5%) |
| All memory | $1.41 (41%) | $0.84 (28%) |
| GPU, excluding HBM | $0.62 (18%) | $0.46 (15%) |
| Scale-up networking | $0.13 (4%) | $0.37 (12%) |
| Scale-out networking | $0.49 (14%) | $0.49 (16%) |
| All-in capital cost | $3.48 (100%) | $3.01 (100%) |
Source: SemiAnalysis, September 1, 2026. Figures are the firm's estimates, not Nvidia disclosures.
The HBM line falls by more than half, from $0.99 to $0.43 per GPU-hour, and SemiAnalysis said that already includes the 2026 HBM price increase. The all-in cost drops 14%, from $3.48 to $3.01. The money doesn't all disappear. "It's redirected into scale-up networking: on the NVL576 NPO SKU, scale-up triples from 4% to 12% of rack spend as optics take over rack-to-rack interconnect," the firm wrote. Scale-up networking is the fabric that links GPUs inside a single system so they work as one machine, and the larger Rubin Ultra design connects 576 of them.



