404K Semi-Ai

404K SEMI-AI Evening Brief, 2026-08-25 — Bottlenecks Spread: AI Rack Costs Flow Through to Memory, Packaging, and End Demand

404K Semi-Ai's avatar
404K Semi-Ai
Aug 25, 2026
∙ Paid

目录

  • Pre-Market Highlights

  • Full AI/Semiconductor Value Chain

  • AI Models, Applications, and Capital Expenditure

  • AI Cloud and Data-Center Operators

  • GPU/CPU/ASIC

  • HBM/DRAM/NAND/SSD/HDD and Advanced Packaging

  • Optical Communications, Cooling, and Passive Components

  • Internet/Platforms

  • Software/SaaS

  • Consumer Electronics / Smart Vehicles

Overview

404K SEMI-AI | 2026-08-25

Pre-Market Highlights

The AI infrastructure bottleneck is shifting from GPU availability to full-system deliverability. Potential price increases of more than 15% for NVIDIA’s new racks reflect simultaneous inflation in HBM, server DRAM, wafers, and cooling, while ABF, advanced packaging, high-end MLCCs, and optical interconnects are spreading the pressure across the supply chain.

Compute architectures are also entering a new phase. Agents are elevating CPU orchestration, long-context memory, and low-latency inference, with AMD, Intel, and Groq all emphasizing rack-level throughput and performance per watt. Cloud providers must now demonstrate not only the scale of capital spending, but also utilization, payback periods, and cash flow.

End demand is clearly diverging. Global smartphone shipments remain under pressure, with entry-level and mid-range devices the first to feel component inflation; Samsung Electronics and Apple are relatively resilient thanks to their supply chains, distribution, and premium portfolios. Humanoid robots and autonomous driving continue to scale, but the next phase will hinge on mass production, real-world reliability, and the pace of cost reduction.

Full AI/Semiconductor Value Chain

AI Models, Applications, and Capital Expenditure

  • OpenAI: Agents are turning model calls into much heavier system workloads. Codex now has more than 20 million weekly active users, while an “automated AI research intern” reportedly runs on a large-scale GPU cluster and targets September 2026 as a capability milestone. Whether demand converts into recurring paid usage will determine whether compute-intensive inference becomes a revenue engine or merely a cost center; revenue per task must cover resource consumption.

"Codex has more than 20 million weekly active users and is testing advertising monetization; Grok Bot drives per-user token consumption to approximately 10x that of conventional chat"

  • Anthropic: Adoption continues to expand across enterprise and scientific use cases. The Claude Science partnership with Eli Lilly is expected to reach an annual run rate of approximately US$1 billion, while Claude Code continues to add CLI, cost-tracking, and organizational model-management features. Model vendors are shifting their value proposition from one-off answers to persistent integration into research and software-development workflows.

  • Alibaba: AI cloud growth continues to require heavier capital investment. June-quarter revenue totaled RMB 269 billion, up 9% YoY, while AI cloud and compute revenue reached RMB 48.4 billion, up 45%. Capital expenditure rose 75% to RMB 67.7 billion, and free cash flow was negative RMB 44.7 billion. The key question is whether the approximately 3-year payback period can be achieved, with revenue growth, returns on capital, and cash collection all requiring validation.

"Alibaba’s June-quarter revenue totaled RMB 269 billion, up 9% YoY, while AI cloud and compute revenue reached RMB 48.4 billion, up 45%"

  • Baidu: Q2 revenue fell 4% to RMB 31.3 billion, while AI cloud infrastructure grew approximately 50% and GPU cloud revenue rose 283%. Legacy-business pressure is coinciding with rapid AI infrastructure growth. The market needs evidence that these high-growth businesses can quickly generate sufficient gross profit and cash flow rather than merely reshuffle the revenue mix; otherwise, earnings quality will continue to deteriorate.

  • Meta: The reliability cost of model training is becoming more visible. Micron disclosed that HBM failures accounted for approximately 17% of unexpected Llama 3 training interruptions. Meta also plans to launch its consumer AI agent Hatch in the coming weeks and the Watermelon model in October. Application expansion must be matched by improvements in infrastructure stability.

AI Cloud and Data-Center Operators

  • Nebius: Token Factory has become the first AI cloud platform to deploy NVIDIA Groq 3 LPX. Each rack supports up to 256 LPUs, 128GB of on-chip SRAM, and 640TB/s of scale-out bandwidth, delivering approximately 3,400 tokens per second on the Gemma 4 31-billion-parameter model. Low-latency inference is emerging as a new vector of cloud differentiation.

  • Lambda: The company is reportedly seeking up to US$3 billion at a valuation above US$12 billion and expects revenue to exceed US$1.5 billion this year. Funding equivalent to nearly 2x revenue highlights how emerging cloud providers must exchange capital for GPU and data-center expansion. Sustaining the valuation will depend on utilization, contract quality, and borrowing costs, while the post-financing balance sheet is another risk factor.

  • Oracle: If NVIDIA raises rack prices by more than 15% in early 2027, cloud providers such as Oracle—which purchase GPUs externally while locking in rental rates through long-term contracts—will face costs rising before lease pricing can reset. Deployment costs could exceed US$60 billion per GW, pushing depreciation and financing expenses more directly into cash flow. Renewal pricing is the clearest validation point.

  • CoreWeave: Rising rack costs pose a greater threat to highly leveraged AI cloud providers. Fixed rental pricing cannot immediately absorb higher HBM, DRAM, and full-rack quotations, while customers may compare AMD or custom-ASIC alternatives. The real defenses are contractual repricing rights, financing duration, and GPU utilization—not simply adding more racks. Cash-payback periods must also keep pace.

  • SpaceX: SpaceX and NVIDIA are reportedly designing a Vera Rubin NVL72 system for orbital environments, targeting a Q4 2027 launch and scaled deployment in 2028. Orbital data centers must simultaneously address power, thermal management, bandwidth, and maintenance. The timeline remains contingent on system-reliability validation, while cost per unit of compute also needs to be quantified.

GPU/CPU/ASIC

  • NVIDIA
    1) Vera Rubin and Grace Blackwell racks could reportedly see price increases of more than 15% from early 2027. The full-rack BOM for VR200 NVL72 is approximately US$7.8 million, versus less than US$4 million for GB300.
    2) Raymond James raised its price target from US$330 to US$352, maintaining a “Strong Buy” rating, and estimates that CPUs could account for 5% of total revenue by 2028.
    3) Rubin power optimization could unlock an additional 18% to 27% of performance at fixed power, though realized gains will depend on facility and system implementation.

"To maintain a 75% gross margin, NVIDIA would need to raise rack prices by 20% to US$24 million"

  • AMD
    1) Helios combines 18 nodes and 72 MI455 GPUs in a single rack, delivering nearly 1.7PB/s of aggregate HBM bandwidth; direct liquid cooling removes more than 85% of the heat.
    2) Raymond James upgraded AMD to “Strong Buy” and raised its price target to US$641, forecasting that the server CPU market will reach approximately US$201 billion by 2030, representing a 44% 5-year CAGR.

"AMD Helios organizes 18 nodes and 72 MI455 GPUs into a single rack-scale system, with nearly 1.7PB/s of aggregate HBM bandwidth"

  • Intel
    1) Diamond Rapids supports up to 256 P-cores, 1.28GB of cache, and 16 memory channels. Crescent Island targets inference with up to 480GB of LPDDR5X, while Wildcat Lake focuses on entry-level PCs and edge applications.
    2) Mercury Research estimates that Intel’s overall x86 share fell to 69.3%, while AMD’s rose to 30.7%. Product updates must translate into share stabilization quickly.

"Intel Diamond Rapids uses a fan-out architecture that separates compute, memory, and I/O, supporting up to 256 P-cores, 1.28GB of cache, 16 memory channels, and approximately 1.6TB/s of peak bandwidth"

  • Groq: Groq 3 LPX has entered full-scale production and is scheduled for deployment at Nebius this year. Each LPU features 500MB of SRAM and 150TB/s of bandwidth, while private endpoints delivered 3,431 tokens per second on the Gemma 4 31-billion-parameter model. Its strength is decoding latency; the risks are ecosystem maturity and validation by large-scale customers.

  • Tensordyne: The company fits 72 chips into one-quarter of a rack using air cooling, with power consumption of 30kW. A comparable Blackwell system could occupy a full rack and consume 150kW. If its logarithmic-mathematics approach can maintain accuracy and software compatibility under real workloads, deployment cost could become a more compelling selling point than peak compute. Customer validation is the next hurdle.

HBM/DRAM/NAND/SSD/HDD and Advanced Packaging

  • Micron: The memory wall presented at Hot Chips continues to widen: AI compute performance increases by up to 3x every 2 years, while HBM bandwidth grows by less than 2x. Eight HBM4 stacks already account for 90% of the semiconductor area in the latest GPU packages, and consume 3x as many wafers as DDR5 DRAM for equivalent capacity. Supply, reliability, and packaging must improve in tandem.

  • Samsung Electronics
    1) Omdia expects DRAM capacity to rise from 7.695 million wafers in 2025 to 8.175 million in 2026 and 8.28 million in 2027, representing growth of approximately 6% and 1.3%, respectively, with the focus shifting to yields on the 1c process.
    2) A related import indicator rose 134% YoY in July. The company guided for Q3 HBM4 sales to increase by more than 3x QoQ and account for over 60% of HBM sales in 2H.

  • SK hynix
    1) Hybrid bonding is now expected to be delayed until HBM5 at the earliest, in 2029–2030, while HBM4E will continue to rely primarily on MR-MUF.
    2) The company plans to adopt molybdenum across its 375-layer ultra-high-stack NAND products. Its next-generation HBM will also reportedly use Intel EMIB, but the near-term priorities remain yields on existing processes and material supply.

  • SanDisk: Data show that SanDisk resumed buybacks in the same year it recorded US$59.8 billion of flash-memory bookings. For every US$1 invested in factories this year, it is spending approximately US$25 on its own shares; the fabs are owned through its joint venture with Kioxia. Capital returns are aggressive, but their sustainability depends on converting bookings into revenue, clearing inventory, and maintaining NAND pricing. Capital allocation between buybacks and capacity expansion also warrants scrutiny.

  • Ajinomoto: The ABF shortage has shifted from nominal capacity to qualified capacity. Ajinomoto holds more than 95% of the global market for films used in high-layer-count AI substrates, and Q2 capacity was fully utilized. The industry expects a shortfall of approximately 8% to 10% in 2026, potentially widening to 35% to 42% by 2028. High-end yields and qualification speed matter more than capacity expansion alone.

"Global sales of interlayer dielectric materials are expected to rise from 19 million square meters in 2025 to 22.9 million in 2026, 27.5 million in 2027, and 41 million square meters in 2030"

User's avatar

Continue reading this post for free, courtesy of 404K Semi-Ai.

Or purchase a paid subscription.
© 2026 lihua · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture