目录
Executive Summary
I. Official Specifications and Unverified Fields
II. Three Levels of Analysis for the HBM Demand Impact
III. Why HBM and DRAM Prices Will Not Immediately Reverse in 2026 as a Result
IV. The Real 2027 Impact Is on the HBM4E Premium and Product Mix
V. Which of the Three Memory Companies Is More Sensitive
VI. Earnings and Valuation Transmit at Different Speeds
VII. What Recent Share Prices Have Already Priced In
VIII. Three Supply-Demand and Valuation Scenarios
Conclusion: The Impact Is Material, but It Affects the Longer-Term Slope, Not Current Pricing
本内容基于公开资料和研报数据整理,不构成任何投资建议,不代表任何个人观点,仅供学习参考,请理性阅读
SemiAnalysis says the mainstream Rubin Ultra configuration has been reduced to HBM4 8-Hi, 192GB, and 1,800W Max-Q. If this supply-chain scenario materializes, the main impact would fall on the 2027 HBM4E product mix and memory-stock valuations, rather than the direction of DRAM prices in 2026.
Executive Summary
SemiAnalysis research dated July 29 says the mainstream Ultra SKU uses HBM4 8-Hi, 192GB, approximately 21TB/s, and 1,800W Max-Q, with the same peak theoretical FLOPs as Rubin and the principal upgrade shifting to a 576-GPU scale-up domain; Nvidia has currently confirmed only the existing Rubin specifications of 288GB HBM4, 12-Hi, and 22TB/s, as well as Ultra NVL576’s 8×72-GPU topology, and has not yet confirmed the aforementioned Ultra specifications.
There are 2 capacity-impact benchmarks: relative to the current Rubin’s 288GB, 192GB represents a 33.3% decline, which a 50% increase in GPU count could offset; relative to the previous 2-die Ultra configuration with 384GB of HBM4E 12-Hi, the decline is 50%, requiring GPU count to double to offset it. The earliest 4-die, 1,024GB configuration also changed the die count, power consumption, and system definition, so the 81.25% mechanical capacity gap cannot be directly equated with a reduction in industry demand.
Based on 192GB per GPU and the official 576-GPU domain, NVL576 contains 110.592TB of HBM, 5.33 times that of Rubin NVL72. A larger scale-up domain does not prove actual shipment volumes, but it indicates that demand is shifting from “high capacity per GPU” toward “more lower-power GPUs”; relative to the previous 384GB Ultra configuration, the number of systems or GPUs would need to increase by 100% to maintain the same bit demand.
The direct impact on 2026 pricing remains limited. Ultra targets 2027, while 2026 demand will be driven by Rubin HBM4, AI ASICs, and server DRAM; inventories at the three major manufacturers are low, contract coverage is increasing, and HBM4 is already ramping. The downside is concentrated in 2027, when the mainstream Ultra configuration may revert from HBM4E to HBM4, affecting incremental demand for higher-stack products, the per-GB premium, and the supplier mix.
The impact on total DRAM is clearly smaller than the impact on the HBM4E mix. If Ultra was originally expected to account for 20% of 2027 HBM bit demand, using the same-architecture benchmark of a reduction from 384GB to 192GB, gross industry-wide HBM bit demand would be revised down by approximately 10%, equivalent to approximately 1.3% of total DRAM bit demand and approximately 3% of theoretical DRAM wafer input. This would reduce the slope of price increases, but would not by itself demonstrate oversupply.
As of July 31, Micron Technology, SK hynix, and Samsung Electronics had declined 32.2%, 41.1%, and 27.6%, respectively, from their highest closing prices of the year, but were still up 188.4%, 163.9%, and 118.9% year to date. The SemiAnalysis scenario is most unfavorable to SK hynix’s forward HBM4E scarcity premium, with Micron Technology in the middle and Samsung Electronics relatively more resilient because of its more diversified businesses and lower catch-up hurdle for 8-Hi; this is only a sensitivity ranking, not a reversal in current-period earnings.
I. Official Specifications and Unverified Fields
Nvidia’s Rubin architecture disclosed on July 21 shows that each Rubin GPU consists of 2 compute dies and provides up to 50 PF of NVFP4 inference performance, 35 PF of NVFP4 training performance, 288GB of HBM4, 12-Hi stacks, and 22TB/s of bandwidth. An NVL72 comprising 72 GPUs contains approximately 20.7TB of HBM in aggregate. Nvidia’s public disclosures on Ultra remain limited to system topology: Rubin Ultra NVL576 consists of 8 MGX NVL racks, each containing 72 Ultra GPUs, forming a 576-GPU NVLink domain.
SemiAnalysis’s July 29 supply-chain research further presents a mainstream Ultra scenario comprising 2 compute dies, HBM4, 8 stacks of 8-Hi, 192GB, approximately 21TB/s, and 1,800W Max-Q, with the same peak theoretical FLOPs as Rubin; there may also be a 2,600—2,800W Max-P version and a 1,200W version for lower-compute workloads such as token decode. Its comparison table uses 35 PF, consistent with Nvidia’s official 35 PF NVFP4 training metric, and therefore cannot be directly compared with the 50 PF sparse-inference metric. The 192GB capacity, power tiers, and NPO rack-to-rack interconnect have not yet appeared in Nvidia’s official Ultra datasheet.
The roadmap’s core change is 1,024GB→768GB→384GB→192GB, accompanied by the cancellation of the 4-die design, a reversion from HBM4E to HBM4, and lower power consumption. The original chart contains ambiguous arithmetic labels for the number of HBM stacks and capacity in the intermediate stages, so stack count is not used in the demand calculation; mainstream capacity, HBM generation, die count, and power consumption are consistent with the accompanying text.
UBS’s April 8 supply-chain report had already concluded that Ultra would retain 2 GPU dies and 8 HBM stacks, with the four-die configuration canceled because of yield and throughput constraints for the ultra-large CoWoS package. For equivalent cloud-computing capacity, the dual-die design delivers a smaller performance uplift and requires more packages, so demand for N3 wafers and HBM may not necessarily decline. JPMorgan’s May 29 global memory model still centered on a 12-Hi/16-Hi HBM4E mix and expected the blended average selling price of HBM to increase by 32% in 2027. SemiAnalysis’s new information extends the downgrade from packaging architecture to the HBM generation, stack height, capacity, and power consumption; the elements of the previous model most in need of adjustment are the blended share of HBM4E and memory content per GPU.
Samsung Electronics’ HBM4E product roadmap released in May provides cross-validation: 12-Hi HBM4E offers 48GB/stack, while the planned 8-Hi and 16-Hi products offer 32GB and 64GB, respectively. If Ultra has 8 HBM stacks and 192GB of total capacity, the average is 24GB/stack, consistent with the 8-Hi HBM4 cited by SemiAnalysis rather than Samsung’s 8-Hi HBM4E. A reversion of the mainstream configuration from HBM4E to HBM4 would affect not only capacity, but also the 2027 high-end product mix and per-GB premium.
II. Three Levels of Analysis for the HBM Demand Impact
The first level is per GPU, but the baselines must be separated. Relative to the current Rubin, a reduction from 288GB to 192GB implies a 33.3% decline in bit demand, which can be offset by a 50% increase in GPU volume; relative to the previous 2-die Ultra HBM4E design, a reduction from 384GB to 192GB implies a 50% decline in bit demand, requiring a 100% increase in GPU volume to offset. The earliest 1,024GB design also changed the compute dies, power consumption, and system topology; the 81.25% mechanical capacity difference reflects only the magnitude of the roadmap change and does not mean HBM industry demand will decline by 81.25%.
The 1,800W Max-Q provides the engineering conditions for a volume offset. Relative to the 2,300W Rubin Max-P, power consumption declines by 21.7%; relative to the 2,600—2,800W Ultra Max-P, it declines by 30.8%—35.7%. Lower power consumption, power-delivery BOM, and cooling requirements can expand the number of deployable GPUs, while 8-Hi also benefits yield and packaging throughput; however, whether these advantages can translate into a doubling of shipments still depends on customer capital expenditure, data-center power availability, and software utilization.
The second level is the system. The definitions of “GPU, compute die, and physical package” are not fully consistent across roadmaps, so system capacity can only be calculated after the unit is explicitly defined.
110.592TB is an unofficial estimate based on Nvidia’s current natural convention of “576 Ultra GPUs” and 192GB per GPU. If SemiAnalysis’s 192GB actually refers to a dual-die physical package while 576 still counts compute dies, total capacity would be 55.296TB; this interpretation is inconsistent with Nvidia’s current official wording and can only be treated as a secondary boundary case. The 144 packages in the earliest 4-die architecture and the current 576 GPUs also do not use the same counting denominator, so the 25% decline from 147.456TB to 110.592TB cannot be directly incorporated into an industry demand model.



