404K SEMI-AI Evening Brief, July 20, 2026 — Inference Capacity Tightens, Memory Price Increases Persist, Software Shifts to Usage-Based Monetization
目录
Pre-Market Takeaways
Full AI/Semiconductor Value Chain
AI Models, Applications, and Capex
CSP/Cloud Capex and AI Cloud
GPU/CPU/ASIC
HBM/DRAM/NAND and Wafer Manufacturing
Advanced Packaging, Testing, Optical Interconnects, and Power
Internet/Platforms
Software/SaaS
Consumer Electronics / Smart Vehicles
Investment-Bank Target Price Changes Over the Past 12 Hours
Lower-cost models have not cooled compute demand: Kimi K3 reached its capacity ceiling less than 48 hours after launch. Memory contract prices, server DRAM spot prices, and NAND demand continue to rise, while hardware bottlenecks are spreading from GPUs to power, substrates, and testing. Software is beginning to diverge: platforms able to charge for usage, context, and security requirements are better positioned to convert AI adoption into revenue.
404K SEMI-AI | 2026-07-20
Pre-Market Takeaways
Korean technology assets underwent sharp deleveraging, with the KOSPI down 4.46% and the KOSDAQ down 5.33%, triggering sell-side circuit breakers in both markets. Pressure came from margin liquidations and the unwinding of crowded positions; there is no evidence yet of widespread AI order cancellations. Whether fundamentals are genuinely weakening should be assessed through hyperscaler capex, server deliveries, and memory contract prices.
Lower model costs and stronger hardware demand are not contradictory. Kimi K3’s lower pricing rapidly pushed demand to the cluster limit, while long context windows, agents, and tiered KV caching are transmitting pressure to GPUs, DRAM, NAND, networking, and power. The strongest near-term evidence comes from long-term data-center contracts, server-rack production, and memory spot premiums. Large rumored server orders unconfirmed by the companies should be treated only as evidence to be disproven.
The dividing line in software is also becoming clearer: seats are no longer the only billing unit. GitLab, MongoDB, and Elastic are seeking to monetize code execution, persistent state reads and writes, and permission indexing, respectively. Security software benefits as agents expand the attack surface. The market will continue rewarding platforms that convert AI usage into renewable revenue while scrutinizing whether contract growth translates into cash flow.
Full AI/Semiconductor Value Chain
AI Models, Applications, and Capex
Moonshot AI/Kimi
1) Kimi K3 has 2.8 trillion parameters and a 1-million-token context window, activating 16 of 896 experts per token; the recommended deployment requires at least 64 accelerators.
2) Less than 48 hours after launch, the company suspended new individual subscriptions as capacity approached its limit, showing that lower-priced models initially amplify usage. The risk is that launch traffic may not translate into long-term retention; actual deployment after weights are released by July 27 will be the key validation point.
“Less than 48 hours after Kimi K3’s launch, new individual subscriptions were suspended because request volumes far exceeded expectations and approached the existing cluster’s capacity limit.”
Anthropic
1) Claude Fable 5 has been removed from the fixed allowance in lower-tier plans, while Max and Team Premium users can access no more than 50% of their allowance; the base allowance has also been reduced by roughly one-third.
2) Pro and Team Standard switch to usage-based billing after a one-time US$100 allowance, with input and output priced at US$10 and US$50 per million tokens, respectively. The pricing changes indicate that high-end inference is being rationed by capacity; the risk is that users migrate to lower-cost models.
Databricks
GPU capacity in Asia is reportedly close to exhaustion, with demand expanding simultaneously across Japan, South Korea, the US, and India, requiring the company to raise financing for additional GPU purchases. This supports continued growth in open-source model hosting and enterprise inference, but procurement volumes and delivery schedules remain unavailable. The key indicators are new financing, the pace of GPU deployment, and actual customer consumption—not reserved demand alone.
OpenAI
The company has formed a small, autonomous “Strategic Futures” team covering catastrophic risks, recursive self-improvement, labor markets, and public policy. This organizational investment shows that frontier-model competition has expanded into governance and policy coordination. From an investment perspective, the relevant validation points remain model-release frequency, enterprise spending, and in-house chip progress; governance discussions cannot substitute for revenue growth.
Alibaba/Qwen
The preview version of Qwen3.8 Max has 2.4 trillion parameters, can be accessed through coding interfaces including Qoder, and is slated for an open-weight release. More capable open models will compress application-layer pricing but also increase hosting, inference, and storage usage. Downloads, cloud token consumption, and enterprise adoption following the weight release must be tracked; parameter count should not be equated directly with commercial market share.
CSP/Cloud Capex and AI Cloud
Google
Expected 2026 capex is approximately US$180–190 billion. The July 22 earnings release will test cloud growth, remaining performance obligations, and AI returns. ASIC shipment estimates also suggest Google’s in-house chip shipments could increase from 2.5 million units in 2025 to 7.1 million in 2028. The larger the capex commitment, the more closely the market will scrutinize the corresponding cloud revenue and free cash flow.
Amazon
AWS Trainium shipments are included in a projected increase in AI ASIC volumes from 4.1 million units in 2025 to 24 million in 2028. In-house chips help reduce inference cost per query while transmitting demand to HBM, CoWoS, substrates, and server DRAM. The risk is that if volume growth merely displaces externally purchased GPUs without creating incremental workloads, the improvement in returns on capital will fall short of expectations.
Microsoft
Microsoft’s Maia is also part of the in-house ASIC expansion cycle, as cloud competition extends beyond GPU procurement into chips, memory configurations, and full-rack design. Aggregate 2026 AI capex across five major cloud providers is estimated at US$725.1 billion, 4.9 times the 2023 level. Microsoft has greater balance-sheet capacity; the key issue is the gap between Azure growth and capex.
“This is not the end of the AI infrastructure cycle, but a validation phase in which profitability must be demonstrated.”
Oracle
Backlog is reported at US$638 billion, with revenue expected to grow at a 31% CAGR through FY2030 as cloud demand continues to accelerate. Another estimate puts Oracle’s total debt-to-EBITDA ratio at approximately 5.6x, well above Microsoft and Google at roughly 0.6x. High order visibility coexists with high leverage; revenue recognition, financing costs, and data-center delivery are the principal validation points.
Hut 8
The company signed a second 15-year, US$9.8 billion net lease at Beacon Point, increasing capacity contracted by the same investment-grade tenant to 704 MW. The 1 GW campus now has a US$19.6 billion base-term contract value and is expected to generate average annual net operating income of US$1.31 billion. Contracted AI data-center capacity has reached 949 MW. Compared with rumored orders, such long-term leases provide more verifiable cash-flow evidence, although delivery and financing remain critical.
“Hut 8 signed a second 15-year, US$9.8 billion lease at Beacon Point, increasing capacity for the same tenant to 704 MW and commercializing the 1 GW campus.”
GPU/CPU/ASIC
Nvidia
Combined GB200/GB300 rack production in June 2026 is estimated at approximately 8,000 units, up 5% month over month. The full-year forecast remains 70,000–80,000 units, more than double the roughly 29,000 units produced last year. Physical deliveries continue to grow, although some Q2 shortfalls at Quanta and Wistron were deferred into the second half. Customers have denied rumors of a roughly US$52 billion GB300 order, which should not be included in demand estimates.
Arm
Jefferies reiterated Buy on July 20 and raised its price target from US$290 to US$320, citing increased AGI CPU orders from new customers and the possibility that SoftBank may launch a GPU. CPUs continue to handle host processing and scheduling across different accelerator architectures, and inference expansion supports the addressable market. The risk is that order visibility and SoftBank’s GPU plans still require product- and customer-level validation.
AMD
The Venice server CPU and MI450/MI455 are reportedly using TSMC’s 2 nm process, while the MI550 is rumored to use four compute dies and advanced 3D packaging. Public code lists Anthropic as a customer with the highest priority score of 30, but this only demonstrates evaluation or testing. Whether the July 22–23 conference provides customer disclosures, performance metrics, supply timelines, and evidence of rack-level software maturity will be critical.
“The highest priority rating suggests Anthropic may be a key customer AMD is pursuing, but the current evidence is more consistent with an evaluation-stage relationship.”
Qualcomm
The company raised its FY2029 non-handset revenue target from US$22 billion to US$40 billion, including US$15 billion from data centers. Its HBC near-memory architecture is scheduled for launch in 2027; Qualcomm claims six times the bandwidth per watt of HBM and a four- to eightfold improvement in decoding efficiency and total cost of ownership. If delivered, the architecture could circumvent some packaging bottlenecks, but customer qualification and the stated 133 TB/s bandwidth still require verification.
Marvell Technology
The company spans custom silicon, switching, optical-electrical conversion, and high-speed interconnects. FY2026 revenue was US$8.195 billion, up 42% year over year, and is expected to reach US$23.26 billion in FY2029. Standalone CXL controllers and the transition to UALink create new entry points. The risk is the long forecast horizon; custom-chip customer concentration, product delivery, and the optical-communications cycle will determine whether the projected growth is realized.
HBM/DRAM/NAND and Wafer Manufacturing
Samsung Electronics
Monthly V-NAND capacity exceeds 100,000 wafers, with approximately 60% shifting to V9 and yields above 80%. The 400-layer-class V10 offers more than 50% higher density than the 286-layer V9. Nvidia’s CMX configuration uses 576 SSDs with total capacity of 9,600 TB, and related NAND demand is expected to rise from 35 million TB this year to more than 100 million TB next year. High prices will encourage capacity expansion; V10 yields and CMX deployment are the key validation points.
SK Hynix
Management expects AI-chip demand to grow 60%–100% in 2027 while describing current memory pricing as “abnormal” and arguing that supply should take priority over margins. Industry estimates indicate that its HBM shipments will increase from 18 billion Gb in 2026 to 24 billion Gb in 2027, with the Yongin Y1 fab scheduled to begin production in February 2027. Additional supply could ease end-market costs but may also compress margins at the cycle peak.
Micron
Meritz reports that spot prices for 64 GB server DRAM have risen above US$3,100–3,400, a premium of approximately 146% to the roughly US$1,380 contract price at the end of June, as sovereign-AI projects and testing demand from new data centers drive urgent purchases. Micron’s Malaysia facility also signed a KRW 1.0668 billion contract for testing equipment. Spot-market transparency is limited; the shortage must be confirmed through actual transaction volumes and subsequent contract prices.
“Server DRAM spot prices have entered a sharp upward trend since mid-July.”
TSMC
The company committed an additional US$100 billion, raising its planned total investment in Arizona to US$265 billion, and increased full-year capex guidance to US$60–64 billion. The first-phase 4 nm facility is already in production, some 5 nm capacity is being converted to 3 nm, and 2 nm is expected to become a new growth driver from Q3. Construction costs in the US are approximately four to five times those in Taiwan, making margin dilution the principal trade-off.

