404K Semi-Ai

404K SEMI-AI Morning Brief, 2026-08-24 — Rising Costs Spread from Compute to Memory, Packaging, and Power

404K Semi-Ai's avatar
404K Semi-Ai
Aug 23, 2026
∙ Paid

目录

  • Post-Close Summary

  • Full AI/Semiconductor Value Chain

  • AI Models, Applications, and Capital Expenditure

  • AI Cloud and Data-Center Operators

  • GPU/CPU/ASIC

  • HBM/DRAM/NAND/SSD/HDD

  • Advanced Packaging and Semiconductor Equipment

  • Space/Satellites

  • Internet/Platforms

  • Software/SaaS

  • Consumer Electronics/Smart Vehicles

  • Jensen Huang’s Selected Portfolio

  • Storage Index Performance

  • U.S.-Listed Storage Stocks’ Share of Trading Turnover

Overview

404K SEMI-AI | 2026-08-24

Post-Close Summary

US equities rose again in the latest session, but gains broadened beyond mega-cap technology stocks. The S&P; 500 gained 0.41%, the Nasdaq 100 0.35%, and the Dow 0.89%, while the equal-weighted S&P; 500 and Russell 2000 advanced 0.63% and 0.77%, respectively. Market breadth improved, with cyclicals and small caps outperforming.

Technology performance remained divergent. Software, cybersecurity, and cloud-computing themes rose 1.43%, 1.50%, and 1.14%, respectively, while semiconductor themes fell 0.40%–0.44%. NVIDIA declined 0.98%, Oracle gained 3.10%, and Tesla rose 5.14%, suggesting that trading interest is broadening from pure chip beta into software, cloud services, and applications.

The more important marginal shift is on the cost side: CoreWeave has raised compute prices, supply-demand gaps in memory and ABF substrates continue to widen, and data-center electricity rates and capacity commitment fees are rising. Demand remains strong. The key test in the next phase is whether cloud providers can pass higher chip, memory, packaging, and power costs on to customers; GPU shipments remain the key demand-side indicator.

Full AI/Semiconductor Value Chain

AI Models, Applications, and Capital Expenditure

  • OpenAI
    1) Sam Altman is positioning the company as a platform, with the product interface and API set to converge gradually; the long-term goal is to bring its capabilities to 100 million new businesses and 8 billion people.
    2) ChatGPT and Codex have been integrated, shifting the product from separate chat and coding tools toward a unified entry point.
    3) The Jalapeño inference chip developed with Broadcom is scheduled to enter data centers by year-end 2026, although training will continue to rely on NVIDIA GPUs.
    4) The model can read and use tens of thousands of pages within seconds; the main validation points are accuracy on long-context tasks, inference cost, and enterprise adoption.

“After that, everything depends on how people use it, build products on top of it, and pursue everything else.”

  • Anthropic: The company has hired Amir Salek, who led Google’s TPU development through its 7th generation, and formed a custom-semiconductor team to design higher-performance, more energy-efficient chips for Claude. It currently retains a multi-chip strategy spanning NVIDIA GPUs, Google TPUs, and Amazon Trainium, indicating that in-house development is intended to expand cost and supply options rather than end external procurement.

  • AI chip-model co-optimization: Model companies are developing chips, while chipmakers are building models. NVIDIA is developing Nemotron 4 and paid $6 billion to license Poolside’s model technology while bringing in approximately 100 engineers; AMD launched Instella, which runs on its own GPUs. Competition is shifting from standalone benchmarks toward joint optimization across models, chips, software, and data centers.

“Performance and cost now depend on how tightly models, chips, software, and data centers are integrated.”

AI Cloud and Data-Center Operators

  • IREN: Horizon 1 is complete, Horizon 2 is expected to be fully operational by the end of September, and Horizons 3 and 4 will follow. Its Microsoft contract is worth $9.7 billion and would generate approximately $1.94 billion in annualized revenue at full deployment. The execution priority is to convert the bottlenecks exposed by Horizon 1 into faster deployment in subsequent phases, rather than focusing solely on contract value.

  • CoreWeave: The company said on its earnings call that it raised prices across compute services by approximately 25% in July, citing the demand environment and improving customer returns on investment. The increase directly confirms that high-end GPU supply remains tight, but it also shifts risk to customers: if reductions in unit inference costs fail to keep pace with rental-price increases, demand could migrate to cheaper neoclouds or self-built clusters.

“Our July pricing adjustment included an approximately 25% increase across SKUs in response to the current demand environment and the improving returns customers are seeing from investments in the CoreWeave platform.”

  • Nebius: Comparable GPUs are priced lower at Nebius and CoreWeave than at Amazon, Microsoft, and Oracle. Newer chips carry higher hourly rates but deliver greater throughput, potentially lowering cost per token. Neoclouds compete on availability and output per dollar; the next test is whether the pricing gap can offset financing, utilization, and customer-concentration risks.

  • AI power demand: Global power demand driven by AI chips is projected to reach approximately 315 GW by 2033, up more than 1,100% from 2025. The US is expected to account for approximately 64% of incremental demand, or about 200 GW. During training, large GPU fleets power up and down simultaneously, causing instantaneous electricity consumption to exceed design capacity by as much as 50%; wear on batteries, generators, and cooling systems will therefore become part of total cost of ownership.

“During model training, hundreds of thousands of GPUs can power on and off simultaneously, causing electricity consumption to spike as much as 50% above design capacity.”

  • Data-center electricity pricing: Tennessee Valley Authority has reportedly approved a dedicated data-center tariff that raises electricity prices by approximately 10%. New projects must also pay a capacity commitment fee of $1.5 million per MW over 3–5 years. Plans simultaneously call for 11–32 GW of new generation capacity, indicating that power access is evolving from an operating expense into an upfront capital constraint.

GPU/CPU/ASIC

  • NVIDIA
    1) A Rubin NVL72 rack is estimated to cost approximately $8 million. A price increase of about 17% could add at least $5 billion to chip costs for a 1 GW data center, with cloud providers likely to pass some of the increase on to customers.
    2) NVIDIA is using neoclouds such as CoreWeave and Nebius to expand GPU distribution and validate new architectures, while assuming greater exposure to customer financing and capacity-absorption risk.
    3) Rubin’s token throughput is expected to reach 10 times that of Blackwell. Whether its lower unit cost can offset the higher rack price will be a key variable in cloud procurement.

“A ~17% price increase could add at least $5 billion to the chip cost of a 1 GW data center, and cloud service providers are expected to pass some of the increase on to customers.”

User's avatar

Continue reading this post for free, courtesy of 404K Semi-Ai.

Or purchase a paid subscription.
© 2026 lihua · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture