Rubin Ultra gets a major spec cut — is even Nvidia feeling the memory price pinch?
- Key Takeaway: Nvidia has adjusted the design of its flagship chip Rubin Ultra in response to surging HBM prices, reducing memory specs while expanding interconnect scale to optimize cost structure. The move has fueled concerns about peaking HBM demand, triggering a sharp sell-off in memory stocks.
- Key Elements:
- Rubin Ultra maintains its peak compute of 35 PFLOPs, but memory is cut to 192GB—lower than Rubin's 288GB—with bandwidth improving by just 1 TB/s, resulting in limited performance gains.
- Cost optimization at the core: HBM's share of total cost drops from nearly 40% to 28%, while interconnect cost share rises from 4% to 12%; per-rack bill of materials falls from roughly $8 million to $6.4 million.
- The spec highlight shifts to "expanded interconnect scale": supporting the NVL576 architecture, up to 576 GPUs can be interconnected to form a super logical GPU—8x the scale of the standard 72-card configuration.
- Market reaction: SK Hynix and Samsung shares both plunged around 8%, while the KOSPI index fell roughly 5%, reflecting investor concerns over AI chipmakers reducing their reliance on HBM.
- The SemiAnalysis report has yet to be confirmed by Nvidia, but its findings have already triggered a reassessment of HBM supply-demand dynamics and pricing power.
Original by Odaily Planet Daily (@OdailyChina)
Author: Azuma (@azuma_eth)

Over the past weekend, an institutional report by well-known investment research firm SemiAnalysis sparked widespread discussion across the entire AI industry.
The core content of the report is that, following SemiAnalysis's late-June disclosure that NVIDIA's original 4-die Rubin Ultra design would be halved, the research firm has now revealed that NVIDIA has provided key customers with a Rubin Ultra preview, but its specifications have declined further from previous expectations.

The Rubin Ultra-related information disclosed in the SemiAnalysis report is as follows:
- Rubin Ultra will maintain the same theoretical peak compute as Rubin, both at 35 PFLOPs;
- Memory capacity has been significantly cut, with specifications reduced to 8-high (8-Hi) stacking at 192GB, even lower than the 12-high (12-Hi) 288GB Rubin;
- Memory bandwidth changes are minimal, with only a 1 TB/s improvement, which is essentially negligible in real-world high-throughput computing;
- Chip-level power consumption has even increased slightly, with minimum power draw matching Rubin (1800W), while maximum power draw is higher (2600W);
- The only major upgrade is the "scale-up world size", increasing from 72 GPUs to 576 GPUs. This means NVIDIA has shifted its key selling point to cluster networking—through NVLink, it can connect up to 576 Rubin Ultra GPUs into a single massive "super logical GPU" (the standard version supports at most 72 interconnected GPUs).
As the flagship top-tier version of Rubin that NVIDIA heavily announced at GTC 2026, Rubin Ultra was originally positioned to handle extreme-scale AI model training and inference demands by integrating more dies and higher-bandwidth memory. However, based on the latest specification details disclosed by SemiAnalysis, NVIDIA has clearly adjusted its design philosophy for Rubin Ultra.
HBM Prices Surging Too Fast, NVIDIA Is Rethinking the Math
Over the past two years, one of the most critical components in the AI industry's expansion has undoubtedly been HBM.
With the explosive demand for AI accelerators such as NVIDIA's H100, H200, and Blackwell, high-bandwidth memory has evolved from a relatively niche high-end storage product into the most constrained link in the entire AI infrastructure. SK Hynix, Samsung, and Micron have all expanded their HBM investments while continuously breaking revenue records, yet the supply-demand imbalance continues to drive HBM prices ever higher.
For AI chip manufacturers, the importance of HBM cannot be overstated—GPUs handle computation while HBM provides high-speed data throughput, and together they determine the efficiency of AI model training and inference. However, the problem is that HBM has become so expensive that it is now affecting the overall economics of AI systems. Taking HBM3 as an example, a single HBM3 module was priced at just $180-220 at the Q2 2025 low, rose to $600-700 in Q1 this year (contract price), and has surged to $700-850 in Q2 (spot price).
According to SemiAnalysis's calculations, as HBM prices have risen, the pure bill-of-materials (BOM) cost for a single Rubin Ultra rack once climbed from approximately $6.6 million to $8 million. After adjusting the design specifications, the cost can be brought back down to approximately $6.4 million.
For NVIDIA, such a massive cost differential raises a very practical question: If the company continues to increase HBM capacity per the original design, will the performance gains justify the costs? Are there better alternatives?
SemiAnalysis provides an answer in its report: the Rubin Ultra adjustments essentially represent NVIDIA re-optimizing the cost structure of AI systems under resource constraints—reducing the relatively expensive HBM configuration (its cost share dropping from nearly 40% to 28%) while redirecting resources toward higher-value scale-up interconnect capabilities (cost share rising from 4% to 12%).
As mentioned earlier, Rubin Ultra's core upgrade direction has now shifted toward "system-level scale-up interconnect capability." The NVL576 architecture supported by Rubin Ultra can connect up to 576 GPUs via NVLink into a unified compute domain, designed to compensate for the adjustments in individual chip specifications through larger-scale system expansion.
Memory Stocks Plunge as Markets Fear Peaking Demand
Likely impacted by this news, memory-related stocks collectively declined at the open of the Korean stock market this morning. As of 11:45 Beijing time, SK Hynix and Samsung—the two HBM leaders—both fell approximately 8%, while the KOSPI index also dropped about 5%.
The market has begun to worry that if the Rubin Ultra specification adjustments revealed by SemiAnalysis are confirmed (NVIDIA has not yet publicly acknowledged the information), it could mean that NVIDIA, as the most pivotal buyer in AI infrastructure, is reducing its demand for HBM.
Over the past two years, as AI expansion demand has grown rapidly, HBM has become the most constrained component in AI infrastructure. Chipmakers including NVIDIA and AMD have continuously increased HBM configurations in their AI accelerators, driving explosive revenue growth for memory manufacturers like SK Hynix, Samsung, and Micron, while also strengthening the latter's pricing power across the entire supply chain.
NVIDIA's choice, however, may signal that AI chip manufacturers are considering reducing individual chips' reliance on high-capacity HBM through optimized hardware design. If this path proves viable, the room for continued price increases by HBM vendors could become significantly constrained.
AI infrastructure construction will inevitably move past the era of "indiscriminate spec-stuffing and mindless price hikes." And now, even NVIDIA—sitting on the iron throne of computing power—has begun to count every penny.


