404K SEMI-AI Evening Brief, July 21, 2026 — Kimi K3 Drives Memory Demand, DRAM Prices Accelerate, Helios Challenges Vera Rubin
目录
Pre-Market Highlights
Full AI/Semiconductor Value Chain
AI Models, Inference Demand, and Capital Expenditure
GPU/CPU/ASIC
HBM/DRAM/NAND/SSD
Foundry, Equipment, and Assembly and Testing
PCBs, Optical Communications, and Power Supply
Internet/Platforms
Software/SaaS
Consumer Electronics / Smart Vehicles
Kimi K3 has shifted the market’s focus from per-inference compute back to model weights, memory capacity, and interconnects. AMD Helios is competing for inference-rack deployments with greater HBM4 capacity, but still trails NVIDIA on price, power consumption, and deployment density. Across storage, spot prices, export unit values, and long-term agreements all indicate that shortages are spilling over from HBM into DDR5 and conventional DRAM.
404K SEMI-AI | 2026-07-21
Pre-Market Highlights
The tension in AI infrastructure is becoming clearer: models can be cheaper and activate fewer parameters per inference, but all weights still require rapid access, while long-context workloads and concurrent agents continue to consume memory. Kimi K3 reached its cluster-capacity ceiling shortly after launch, suggesting that lower prices are more likely to amplify usage before reducing hardware demand.
Hardware competition is shifting from individual accelerators to full-rack economics. AMD Helios offers more CPU cores and greater HBM4 capacity, making it well suited to inference for trillion-parameter models. NVIDIA’s Vera Rubin NVL72 retains advantages in price, interconnect latency, software, and data-center density. The next points to watch are cost per token and deployment at scale, not headline specifications alone.
Storage provides today’s strongest cyclical signal. South Korea’s DRAM export unit value rose sharply in the first 20 days of July, and Morgan Stanley expects like-for-like prices to increase at least 25% QoQ in Q3. The risk is equally direct: higher prices raise server, PC, and smartphone bill-of-materials costs, which must ultimately be absorbed through downstream price increases, product upgrades, or demand contraction.
Full AI/Semiconductor Value Chain
AI Models, Inference Demand, and Capital Expenditure
Moonshot AI
1) Kimi K3 was released on July 16 with 2.8 trillion total parameters, activating 16 of 896 experts per token and supporting a one-million-token context window.
2) The theoretical lower bound for its 4-bit weights is approximately 1.4TB. The existing cluster approached its capacity limit within 48 hours of launch, prompting the company to suspend new subscriptions on July 19. Sparse activation reduces compute per inference, but low-latency access to all weights continues to burden HBM, DRAM, and networking.
“Compute cost depends on the 16 administrators currently working; memory cost depends on the size of the entire room.”
Inference infrastructure: AI demand is shifting from humans directly using chatbots to agents invoking other agents. Longer contexts, more concurrent tasks, and more complex reasoning simultaneously increase demand for GPUs, HBM, networking, and power. The key metrics are tokens generated by each model, enterprise-agent concurrency, and whether price reductions drive sufficient growth in usage.
“What matters is not the growing number of models, but the exponential growth in inference tokens generated by each model.”
GPU/CPU/ASIC
NVIDIA
1) Vera Rubin NVL72 comprises 36 Vera CPUs and 72 Rubin GPUs, with 20.736TB of rack-level HBM4 capacity and approximately 260TB/s of intra-rack interconnect bandwidth.
2) Each rack is quoted at approximately $3.5 million to $4.0 million, consumes 190–230kW, and weighs approximately 1,600–1,814kg. NVIDIA’s advantages remain its lower rack-level price, the low latency of NVLink 6, CUDA 13 and TensorRT-LLM, and higher data-center deployment density.
“However, NVLink still retains a modest advantage in transmission latency.”
AMD
1) Helios uses 18 Zen 6 Venice CPUs and 72 MI455X accelerators. Its CPUs provide a combined 4,608 cores, nearly 45% more than NVL72, while rack-level HBM4 capacity reaches 31.104TB, 50% more than its competitor.
2) Microsoft has procured Helios and integrated it into Azure AI; OpenAI and Meta previously signed compute orders of 6GW each. Risks include an estimated selling price of $5.0 million to $5.5 million, peak power consumption of 245kW, and a weight of approximately 3,175kg. Large-scale deployment must demonstrate that cost per token can offset these disadvantages.
“Otherwise, this price is entirely uncompetitive.”
Intel
1) Morgan Stanley raised its price target from $73 to $75 but remains concerned about limited commitments from large foundry customers, Intel’s execution track record, and competition from AMD, Google Axion, and Amazon Graviton.
2) Positives include improving server operations, CPU pricing, AI-agent workloads, and data-center demand. The 18A and 14A nodes are targeting risk production in 2028 and volume production in 2029, respectively. Separately, Data Center and AI Group revenue reportedly increased 22% YoY in 1Q26, while the company continues to assess layoffs.
Broadcom: The CPO supply-chain map includes Broadcom across photonic integrated-circuit design, ASIC and xPU design, and lasers, indicating revenue opportunities spanning compute, switching, and optoelectronic components. If open models such as Kimi K3 distribute inference across more cloud and enterprise endpoints, the key validation points for custom ASICs and optical interconnects will be volume-production customers, port-speed upgrades, and system-level power consumption.
Marvell Technology
1) The CPO supply-chain map covers Marvell’s roles in photonic integrated circuits, ASICs and xPUs, retimers, SerDes, and PHYs.
2) One engineer said demand for 800G optical modules exceeds 80 million units, versus capacity of approximately 60 million. Demand for 1.6T modules could exceed 30 million units, with deliveries slightly above 20 million. The shortage is a genuine catalyst, although the engineer also expects module prices not to rise every year.
“The highest-volume product shipped this year is the 800G optical module. Demand exceeds 80 million units, but only about 60 million can be produced.”
HBM/DRAM/NAND/SSD
Micron
1) Bank of America rates Micron Buy with a $1,550 price target, estimating that each Kimi K3 serving instance still requires approximately 1.4TB of high-bandwidth memory and that self-hosting open weights will increase the number of memory sockets globally.
2) A storage expert expects Micron to maintain an HBM share of approximately 20% and benefit from supply shortages. The larger risk is downstream cost tolerance: if elevated prices suppress PC, smartphone, or server demand, spot-market strength may not fully translate into long-term profits.
SK hynix
1) A storage expert estimates its HBM share at approximately 58%–60%, ahead of Samsung Electronics and Micron. The key issue for next-generation products is whether it can integrate logic dies.
2) While fulfilling HBM commitments, SK hynix and Samsung Electronics are redirecting flexible DRAM capacity toward server DDR5. Continued price increases in Q3 support profitability, but competition after 2027 will depend on HBM4, logic dies, and the pace of new fab capacity.
Samsung Electronics
1) An expert estimates its DRAM share at approximately 40% and HBM share at approximately 20%. Its in-house foundry capabilities could help it recapture some HBM share in 2027–2028.
2) With DDR5 margins approaching those of HBM, the company is allocating additional DRAM capacity to server products. Shareholders have objected to special performance bonuses, reminding the market that capital allocation and shareholder returns still require scrutiny during an industry upcycle.
“Memory supply is severely inadequate relative to the level of AI demand, and we do not expect this trend to change.”
SanDisk: As open-weight models are downloaded to more enterprise, cloud, and sovereign endpoints, demand will extend beyond HBM into NAND and storage devices. South Korea exported $969 million of flash memory in the first 20 days of July, up 236% YoY but down 21% MoM. Strong pricing and monthly shipment volatility coexist; the next question is whether enterprise SSDs and high-capacity NAND can absorb growth in inference data.
Kioxia: Kioxia and SanDisk are both positioned in the NAND-capacity expansion driven by open models, but near-term data are not uniformly positive. SSD exports totaled $1.769 billion in the first 20 days of July, up 359% YoY but down 36% MoM. Demand-side indicators include growth in self-hosted endpoints and datasets, while supply-side metrics include NAND capital expenditure, contract pricing, and the enterprise-product mix.
Memory pricing
1) South Korean DRAM exports excluding modules totaled $8.774 billion in the first 20 days of July, up 453% YoY and 18% MoM. Export unit value reached $100,198/kg, up 527% YoY and 22% MoM.
2) Spot prices are trading at a premium of more than 140% to June contract prices. Morgan Stanley expects like-for-like product prices to increase at least 25% QoQ in Q3. If urgent-order pricing continues to lead contract pricing, memory manufacturers will gain stronger pricing power, while downstream bill-of-materials costs will rise more rapidly.

