Nvidia (NVDA) cut Rubin Ultra's memory in half because HBM had become 40% of a rack's cost, SemiAnalysis says

Nvidia (NVDA) cut Rubin Ultra's memory in half because HBM had become 40% of a rack's cost, SemiAnalysis says

Key points

  • Memory had reached 41% of a Rubin Ultra rack's capital cost
  • Halving HBM per GPU takes that to 28%, SemiAnalysis estimates
  • The savings move to scale-up networking, 4% to 12%
  • SemiAnalysis calls it a supply-crunch symptom, not weak demand

SemiAnalysis put a number on why Nvidia (NVDA) stripped memory out of its next flagship AI system. After this year's price increases for high-bandwidth memory and DRAM, memory had grown to about 41% of the total capital cost of a Rubin Ultra rack, the research firm said in a post on X at 9 p.m. ET on Tuesday, September 1, after the market closed. Cutting the memory per chip in half brings that share down to 28%.

The change itself isn't new. The Information reported on August 5 that Nvidia was testing versions of Rubin Ultra with less memory than the 1 terabyte per package Jensen Huang described at GTC in 2025, and SK Hynix lost about 15% in Seoul over the two sessions that followed. SemiAnalysis itself said on June 30 that Nvidia had cut Rubin Ultra from four compute dies to two. The configuration SemiAnalysis now treats as settled uses HBM4 stacked eight dies high, for 192 gigabytes per GPU, in place of the HBM4E 12-high stacks that would have carried 384 gigabytes. That's below the 288 gigabytes on the regular Rubin chip that comes first.

What's new is the cost math, which the firm said it presented at a recent conference keynote. "Why strip memory out of your flagship rack-scale system? Because after the latest HBM and DRAM price hikes, memory had quietly become ~40% of total capital cost of ownership," the post said. The table SemiAnalysis published compares the two designs on a cost-per-GPU-hour basis and assumes a 70% markup on HBM and a 60% markup on DRAM for the GPU provider.

Cost layerRubin Ultra HBM4E 12-Hi (NVL72)Rubin Ultra HBM4 8-Hi (NVL572)
HBM$0.99 per GPU-hour (29%)$0.43 (14%)
DRAM$0.26 (7%)$0.26 (9%)
NAND$0.16 (5%)$0.16 (5%)
All memory$1.41 (41%)$0.84 (28%)
GPU, excluding HBM$0.62 (18%)$0.46 (15%)
Scale-up networking$0.13 (4%)$0.37 (12%)
Scale-out networking$0.49 (14%)$0.49 (16%)
All-in capital cost$3.48 (100%)$3.01 (100%)

Source: SemiAnalysis, September 1, 2026. Figures are the firm's estimates, not Nvidia disclosures.

The HBM line falls by more than half, from $0.99 to $0.43 per GPU-hour, and SemiAnalysis said that already includes the 2026 HBM price increase. The all-in cost drops 14%, from $3.48 to $3.01. The money doesn't all disappear. "It's redirected into scale-up networking: on the NVL576 NPO SKU, scale-up triples from 4% to 12% of rack spend as optics take over rack-to-rack interconnect," the firm wrote. Scale-up networking is the fabric that links GPUs inside a single system so they work as one machine, and the larger Rubin Ultra design connects 576 of them.

Advertisement

Who gets the pricing power

SemiAnalysis read the cut as a sign of strength for the memory makers, not weakness. "When the industry's largest, best-positioned buyer is stripping memory content just to manage cost and availability, pricing power sits firmly with suppliers — and the shortage is broadening, not easing," the post said. That's the opposite of how Korean investors traded the news in August, when SK Hynix (SKHY) fell on the idea that a smaller Nvidia order meant less revenue.

The price data from the past month sits on SemiAnalysis's side. Korea's average HBM export price reached $76.13 per chip at the end of July, up 9.5% in a month, and spot HBM3E was selling near $2,100 a chip on August 31, four to five times the long-term contract rate. Cantor and Fubon Research have each estimated SK Hynix's HBM4 at $31 to $32 per gigabyte for Nvidia, close to double the HBM3E price. Elon Musk said on August 5 that memory demand is growing about 200% a year against about 20% more supply.

The thread went up after Tuesday's close. By 11:30 a.m. ET on Wednesday, September 2, Nvidia was trading at $226.84, up 4.3% from Tuesday's close of $217.44. Micron (MU) was at $940.22, up 0.7%, and SK Hynix's US shares were at $161.65, up 0.5%. Broadcom (AVGO), which reports earnings after Wednesday's close, was flat at $370.00. Astera Labs (ALAB) was down 3.8% at $269.22, and Credo (CRDO) was down 19% at $167.79 after its own first-quarter report on Tuesday night showed gross margin falling to 64.5% from 68.2% the prior quarter. Nvidia hasn't commented on the Rubin Ultra memory configuration. The company said in May that Vera Rubin, the generation before Rubin Ultra, is in full production.

Advertisement

Frequently asked questions

Why did Nvidia cut the memory on Rubin Ultra?

According to SemiAnalysis, memory had become about 41% of the total capital cost of a Rubin Ultra rack after this year's HBM and DRAM price increases. Moving from HBM4E 12-high stacks (384GB per GPU) to HBM4 8-high stacks (192GB per GPU) cuts the HBM cost by more than half and brings memory down to about 28% of the rack's cost. Nvidia has not commented on the configuration; the figures are SemiAnalysis estimates presented at a conference keynote and posted on X on September 1, 2026.

How much does the change save per rack?

SemiAnalysis estimates the all-in capital cost falls from $3.48 to $3.01 per GPU-hour, about 14%. The HBM line drops from $0.99 to $0.43 per GPU-hour. Part of the savings is redirected to scale-up networking, which rises from 4% to 12% of rack spend on the NVL576 configuration as optics take over rack-to-rack interconnect.

Is the Rubin Ultra memory cut bad news for SK Hynix (SKHY), Samsung, and Micron (MU)?

SemiAnalysis argues the opposite: when the largest buyer strips memory content to manage cost and availability, pricing power sits with the suppliers and the shortage is broadening. Korean investors sold SK Hynix about 15% over two sessions in early August on the first reports of the cut. Korea's average HBM export price rose 9.5% in July to $76.13 per chip, and spot HBM3E was selling at four to five times the contract price at the end of August. This is general information, not investment advice.

Which companies benefit from the shift to scale-up networking?

Scale-up networking links the GPUs inside a single system, and SemiAnalysis says it triples as a share of rack spend under the new Rubin Ultra design. Companies that sell interconnect and optical parts for that fabric include Broadcom (AVGO), Astera Labs (ALAB), and Credo (CRDO). Nvidia's own NVLink and co-packaged optics switches also sit in that layer. This is general information, not investment advice.

More on NVDA and MU

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.