目录
Pre-Market Highlights
Full AI/Semiconductor Value Chain
AI Models/Applications and Capital Expenditure
GPU/CPU/ASIC
HBM/DRAM/NAND/SSD/HDD
Foundries, Equipment, and PCBs
Optical Communications/Optical Supply Chain
Servers, Networking, and Power
Internet/Platforms
Software/SaaS
Consumer Electronics/Smart Vehicles
HBM and indium phosphide lasers are simultaneously becoming constraints on AI infrastructure expansion, while Google is shifting more TPUs toward external sales and leasing. Hardware demand remains strong, but profits are becoming more dependent on supply, yields, and contracted capacity.
404K SEMI-AI | 2026-08-07
Pre-Market Highlights
The key tension in AI infrastructure continues to shift from “is there demand?” to “who can deliver on time?” HBM price expectations are moving higher, and any reduction in HBM for Rubin Ultra would sacrifice throughput and compute utilization; the optical chain, meanwhile, faces multiple constraints across EMLs, indium phosphide substrates, lasers, and manufacturing capacity. Both chains point to the same conclusion: the higher the compute density, the harder it is to economize on memory and interconnects, and the more concentrated the constraints become.
Another divergence is emerging in the cloud. The competitiveness of Google’s frontier models is being questioned, but external TPU sales, leasing, and complete-system deliveries give GCP a more direct path to revenue and profit. Akamai’s sold-out GPU capacity, long-term capacity-lock contracts, and high cash gross margins also indicate that new-cloud demand has not yet entered a phase of material price declines.
For software and end devices, the focus is increasingly on product execution. Cloudflare, JFrog, and Innodata have provided growth or order signals, while OpenAI’s hardware remains at the stage of clues about pricing, form factor, and launch timing. Whether consumers can absorb higher storage and inference costs will still depend on actual sales, retention, and unit economics.
Full AI/Semiconductor Value Chain
AI Models/Applications and Capital Expenditure
Google
1) Gemini API tokens per minute rose from 10 billion to 16 billion in Q1 2026, up 60% quarter over quarter; they increased to 22 billion in Q2, with growth slowing to 38%.
2) Google is commercializing more TPUs: from Q3 2026 through Q4 2027, more than 20% of TPU shipments are expected to be sold directly to Anthropic. GCP growth is beginning to decouple from Gemini’s competitiveness; the next key indicators are the pace at which TPU orders convert into revenue and RPO.
“GCP’s third-party AI IaaS/TaaS ARR will exceed $73 billion.”
SoftBank Group: Borrowed $10 billion against its OpenAI stake in a loan led by Goldman Sachs and JPMorgan at an interest rate of 7.88%; SoftBank Group has committed to investing $65 billion in OpenAI and has total debt of $135 billion. The financing cost is significantly higher than for conventional loans secured by publicly traded equities, indicating that AI capital formation continues to expand, while the lack of public liquidity for private equity is also raising funding costs.
“Because OpenAI is a private company and its shares are not traded, banks cannot independently assess the value of the collateral, so SoftBank Group had to provide a corporate guarantee.”
Anthropic: After Claude Fable 5 adjusted its biological safety safeguards, related deferrals in testing declined by approximately 85%, allowing it to handle more routine health and education questions; dual-use requests involving areas such as virology, toxicology, and molecular design are still referred to Opus 5. The product’s usable scope has expanded, but its professional research capabilities remain constrained by safety boundaries.
“For requests we consider dual-use, including virology, toxicology, and molecular design, Fable will continue to defer to Opus 5.”
GPU/CPU/ASIC
GPU rental pricing: The price of a one-year H100 contract rose from $1.70 per GPU-hour in October 2025 to $2.60, an increase of nearly 40%; B200 quotes range from $4.99 to $18 per GPU-hour depending on the provider. Lambda and Verda have also raised their published rates. Prices have not yet shown the declines typically associated with oversupply; the real counterevidence would still be a slowdown in laboratory usage or a sustained collapse in GPU prices.
NVIDIA
1) A fully configured Rubin Ultra with 16 HBM4E stacks provides 768GB of memory and approximately 64 TB/s of bandwidth; a 12-HBM4E configuration would provide approximately 50 TB/s, with an estimated performance loss of 20% to 25%, making it the more viable fallback under tight supply.
2) If bandwidth falls from approximately 60 TB/s to 20 to 30 TB/s, large-language-model inference throughput may decline by 40% to 60%. Saving on HBM would translate into more parallel communication, latency, and idle GPU time, while making the flagship premium more difficult to sustain.
3) NVIDIA’s software ecosystem: Cooperative Matrix 2 reportedly brings Vulkan to approximately 95% of CUDA’s level; by comparison, AMD’s ROCm still has unresolved issues. Beyond peak hardware performance, compilers, toolchains, and the developer experience will continue to affect customers’ actual utilization rates, making this an important test of how quickly alternative platforms can catch up.
“16 HBM4E stacks are a mandatory configuration for fully maintaining the Ultra positioning; 12 HBM4E stacks are the most viable alternative when balancing performance, cost, and topology.”
AMD: AMD is reportedly acquiring Taalas, which hardwires AI models into custom chips. If the transaction is completed, AMD would gain not only chip designs but also a team and methodology for model-specific optimization; the next questions are whether the technology will be integrated into AMD’s existing GPU and software roadmaps and whether it can produce a repeatably deliverable product.
South Korean AI chip startups: Independent NPU vendors face two simultaneous pressures: Google, Amazon, and Meta can plan ASICs 2 to 3 years in advance around their own service roadmaps, while NVIDIA continues to dominate the general-purpose commercial chip market. Rebellions’ acquisition of model-compression company SqueezeBits to strengthen its software capabilities indicates that competition has expanded from single-chip performance to tools, optimization, and customer support.
Akamai
1) All existing GPU capacity is sold out, with customers tending to lock in capacity over the long term and reserve several months in advance; large cloud transactions use take-or-pay commitments, with cash gross margins of approximately 65% to 75% and operating margins of approximately 20% to 30%.
2) The company also disclosed a 4-year, $600 million GPU services contract, bringing the total value of major orders announced recently in 2026 to more than $2.8 billion. Supply and demand remain tight; the next question is whether the project pipeline can be executed at the margins stipulated in signed contracts.
“GPU demand remains exceptionally strong, and all of our GPU capacity is completely sold out.”
HBM/DRAM/NAND/SSD/HDD
HBM industry: Goldman Sachs expects average HBM prices to remain at $1.4 to $1.7 per gigabit in 2025 to 2026, rise to approximately $2.9 in 2027, and reach $3.0 in 2028; Samsung Electronics’ and SK hynix’s average HBM prices may rise 87% to 100% year over year in 2027. Another industry view expects memory capacity to grow only 20% to 30% annually over the next 3 years, while demand may double. The transition to HBM4, higher stack counts, and yield challenges are collectively reinforcing suppliers’ pricing power.
“This reflects a structurally tight supply environment, rather than a short-term price spike.”
Micron
1) Following discussions with management, Deutsche Bank noted that both DRAM and NAND are in shortage, while memory’s share of total system value has risen from approximately 10% 30 years ago to nearly 50%.
2) Strategic customer agreements are expected to cover approximately 40% of sales and include volume, duration, and price floors and ceilings; Micron’s disclosed minimum contracted revenue has exceeded $100 billion. These contracts improve revenue visibility, but views remain divided on whether the market will award a higher valuation.
“Memory’s share of total system value has risen to nearly 50%, and AI is accelerating this revaluation process.”
SK hynix
1) The company plans to add approximately KRW54 trillion in fab investments through 2031, including approximately $24.9 billion for Yongin Phase 2 and approximately $13.5 billion for Cheongju M17, while bringing forward completion of the Yongin cluster from 2045 to 2033.
2) The company denied supplying HBM to NVIDIA at 50% below market price, emphasizing that specifications, qualification, volumes, and contract durations differ by customer and that HBM has no uniform market price that permits simple comparisons.
Samsung Electronics: Management said it would not allocate more than 60% to 70% of memory capacity to multiyear contracts, not because of a lack of offtakers but because it needs to retain flexibility across spot sales and customer mix. Nearly all customers are requesting long-term agreements, indicating that cyclical products are acquiring more backlog-like characteristics; however, the cap on contracted volumes also reminds the market that long-term agreements do not mean all capacity and pricing are fixed.

