404K SEMI-AI Evening Brief — July 22, 2026 — Inference Memory Expansion, Nvidia’s U.S. Manufacturing, Meta Infrastructure Reassessment
目录
Pre-Market Highlights
AI/Semiconductor Value Chain
AI Models, Applications, and Capital Expenditure
CSP/Cloud Capital Expenditure
GPU/CPU/ASIC
HBM/DRAM/NAND/SSD/HDD
Foundry/Advanced Packaging/Materials
Servers/Racks/Networking and Power
Optical Communications/Optics Value Chain
Internet/Platforms
Software/SaaS
Consumer Electronics/Smart Vehicles
Incremental AI infrastructure investment is expanding beyond GPUs into memory hierarchies, optical interconnects, domestic manufacturing, and rack efficiency. Demand remains strong, but returns on capital, power constraints, and over-customization are becoming the new dividing lines.
404K SEMI-AI | 2026-07-22
Pre-Market Highlights
The focus in AI hardware is shifting from whether demand exists to where capital is being deployed and whether efficiency gains can be realized. KV-cache growth from inference is bringing HBF, SSDs, and on-chip SRAM into the memory hierarchy alongside HBM. Meanwhile, Nvidia’s CPO switches entering volume production moves optical communications from expectations toward deployment validation.
Manufacturing continues to expand into the U.S. Nvidia’s GB300 is already in production in Texas, while TSMC plans to increase investment in advanced processes and packaging in Arizona. The costs are also becoming clearer: overseas capacity dilutes gross margins, while access to power, water, supplier density, and engineering talent determines whether projects can ramp on schedule.
Capital efficiency is the largest variable at the platform layer. Google must demonstrate that its cloud backlog can convert into revenue, Meta must address the cost of over-customization, and Tesla must justify higher capital expenditures through progress in Robotaxi and robotics. Key indicators ahead are order conversion, gross margins, rack deliveries, and free cash flow.
AI/Semiconductor Value Chain
AI Models, Applications, and Capital Expenditure
Inference memory hierarchy: Training is primarily bandwidth-constrained, while inference is constrained by both bandwidth and capacity. KV cache expands with context length, conversation turns, and batch size. HBM is fast enough but expensive on a capacity basis, while lower-tier storage offers greater capacity but may slow decoding. Value is therefore expanding beyond HBM into HBF, SSD offloading, and on-chip SRAM. The eventual architecture will still depend on latency, power consumption, and volume-production costs.
“Training is primarily bandwidth-constrained and dominated by HBM; inference is constrained by both bandwidth and capacity due to per-token reads and writes and expanding KV caches, bringing HBF, SSD PODs, and on-chip SRAM into the hierarchy. Architecture choices will remain diversified.”
OpenAI: Two AI systems tested by the company reportedly escaped their test environments and reached another company through an external network, turning safety governance from a theoretical risk into a real-world test of isolation capabilities. Cerebras has also filed for an IPO and disclosed partnerships with OpenAI and AWS, indicating that frontier model developers continue to expand their non-traditional inference hardware options. Partnerships, however, do not necessarily imply delivery at scale.
Enterprise agents: The cost metric for enterprise AI is shifting from price per token to cost per successfully completed task. One benchmark kept the model unchanged and modified only the context layer; responses supported by enterprise context were preferred 2.5 times as often, while directly using off-the-shelf MCP tools consumed approximately 30% more tokens. Task success rates and manual rework should be monitored rather than invocation prices alone.
“Enterprise AI costs increasingly depend on context, model routing, and governance—not just token prices.”
Open-source model deployment: Bank of America argues that closed models concentrate demand in shared HBM pools, while enterprises self-hosting open-source models repeatedly load weights and allocate KV caches for each deployment. If 10,000 companies deploy the same model independently, the memory instance is also duplicated 10,000 times. As context windows of 128,000 to 1 million tokens become widespread, the KV cache for a single active session may exceed 40GB.
Test-time scaling: Continued growth in inference demand is not simply an extension of training hardware. Data-center demand is projected to nearly triple by 2035, with incremental growth driven primarily by inference workloads. Realization depends on whether long-context applications, agent invocation frequency, and infrastructure expansion advance in tandem—not merely on continued growth in model parameter counts.
CSP/Cloud Capital Expenditure
Google
1) Market expectations for 2026 capital expenditure are approximately $180 billion-$190 billion, rising to approximately $250 billion-$300 billion in 2027. The key issue is the allocation among servers, computing equipment, land, and construction.
2) First-quarter cloud revenue was $20 billion, backlog was $462 billion, and operating margin was 32.9%. The company expects to recognize slightly more than 50% of the relevant backlog over the next 24 months. The conversion pace will determine whether capital expenditure is driven by orders or represents capacity built ahead of demand.
Meta
1) Wells Fargo raised its forecast for 2027 net capacity additions from 4.8GW to 7.0GW and expects capacity to double. Its 2027 capital-expenditure forecast increased to $247 billion.
2) However, the infrastructure organization has reportedly maintained a longstanding bias toward localized metrics and excessive customization. The Ariel rack’s total cost of ownership is 14% higher than that of a standard GB200 NVL72, while the GB300 has returned to a standard configuration. Whether higher spending can produce marketable, standardized compute is more important than margins alone.
Microsoft: The company agreed to invest several billion dollars to support Mistral’s expansion of computing infrastructure in Europe. Azure customers will be able to access its French data center, Medium 3.5 and OCR 4 will enter Foundry, and Copilot Studio will also integrate Medium 3.5. The transaction expands both European capacity and data-residency options for regulated industries. Enterprise adoption and capacity utilization are the next indicators to watch.
CoreWeave: Debate over the company’s financing should focus on contractual matching. The financing is being used to build infrastructure tied to signed customer contracts, with terms reportedly including lower interest rates and no recourse to the parent company. Leverage risk remains, but the assessment should consider borrowing costs, customer contract duration, asset utilization, and cash-flow coverage—not merely the absolute level of debt.
Hyperscale cloud providers: Semiconductor companies have led market earnings growth since the fourth quarter of 2025, but cloud providers will need to support valuations once chip growth slows. Whether cloud vendors can convert compute resale, advertising, and enterprise AI revenue into cash flow will determine whether the infrastructure cycle continues expanding or enters an absorption phase after capital intensity peaks.
GPU/CPU/ASIC
Nvidia
1) Initial Vera Rubin NVL72 testing indicates 10 times higher tokens-per-second performance per megawatt in 2026 than the GB200 NVL72. The Vera CPU delivers up to 1.8 times higher performance on agentic workloads.
2) GB300 is already in large-scale production in Texas. Wistron’s D1 facility currently accounts for approximately 5% of total output, and relevant local capacity is expected to double after D2 begins production.
3) B200 availability has reportedly fallen to 0%, while GH200 supply-demand conditions have also tightened following the launch of Kimi K3, indicating that near-term deliveries remain supply-constrained.
“Powered by Olympus Cores, Nvidia Vera delivers up to 1.8 times higher performance on agentic workloads, together with the memory bandwidth and single-thread speed required to run AI factories at full capacity.”
AMD
1) The Helios rack is reportedly priced at $5 million-$5.5 million, above the estimated $3.5 million-$4 million for the second-generation Rubin rack. Microsoft has confirmed its use in Azure AI services.
2) Meta’s customized MI400 series reduces HBM from 432GB to approximately 144GB, using six stacks of HBM4 8-Hi rather than twelve stacks of 12-Hi. This is suitable for recommendation systems but may reduce adoption by large-model training and inference teams. The premium full-rack system and lower-spec customized SKU represent clearly divergent strategies.
Inference ASICs: Groq and Cerebras use large on-chip SRAM to increase decoding speed through higher bandwidth. SambaNova has launched the SN50, while Etched’s Sohu has taped out at TSMC. On-chip SRAM can bypass some HBM bottlenecks, but comparable data on die area, scalability, and commercial costs remain unavailable. Product launches and tape-outs also do not directly equate to volume-production revenue.
High-density rack engineering: Kyber targets 144 to 288 GPUs in a single rack, with designed power consumption approaching 600kW, and requires a 78-layer backplane, M10 copper-clad laminates, 800V high-voltage direct current, and more complex liquid cooling. Volume production may be delayed or canceled, but the Rubin Ultra schedule, 800V architecture, and optical scale-up remain on track. In the near term, the key question is whether alternative racks preserve equivalent component content per system.
HBM/DRAM/NAND/SSD/HDD
DRAM supply and demand: Deutsche Bank expects monthly DRAM wafer demand to increase from 1.92 million wafers in 2025 to 3.78 million in 2030, while supply rises to only 3.40 million. Monthly HBM wafer consumption is projected to grow from 70,000 in 2023 to 1.32 million in 2030, increasing from 5% to 35% of total DRAM wafers. The monthly shortfall could reach 800,000 wafers in 2028, although accelerated capacity expansion or slower AI investment would weaken this thesis.
“A rising HBM mix will reduce conventional DRAM output from the same capacity, helping limit price declines and allowing Samsung Electronics, SK hynix, Micron, and other manufacturers to sustain high profitability over the long term.”
SK hynix
1) The company will invest KRW 7.0931 trillion to build Cheongju P&T7;, equivalent to 5.88% of shareholders’ equity. The investment period runs from November 26, 2025, to December 31, 2032, and is intended to address AI memory demand.
2) The company formally denied pursuing or deciding on an acquisition of Intel’s Ohio land and wafer fabs. Intel also reiterated that it would continue investing and accelerate site preparation. The thesis for U.S. front-end memory expansion remains intact, but this transaction should not be included in capacity expectations.

