AI Agent Hardware Relay: CPU-to-GPU Ratio Approaches 1:1 as Rising Memory Prices Crowd Out Consumer Electronics
目录
TL;DR
The Real New Insight in This Monthly Report Is Not a Larger CPU Market Estimate
From 1:8 to Nearly 1:1: The Change Is in the Workload, Not the Chip Narrative
The Simultaneous Rise of Both Server Curves Is the Core of the “Hardware Relay”
Monthly Data Are Turning the Architecture Narrative into Revenue
Memory Prices Are the Most Sensitive Inflection Point in This Cycle
AI Servers Accept the Same Price Increases That Consumer Electronics Reject
Why South Korea’s Massive Capacity Expansion Will Not Immediately Create Oversupply
The Value Chain Is No Longer Limited to GPUs, With Profits Spreading in 5 Directions
From Revenue to Profit: Who Benefits from Volume, Who Benefits from Pricing, and Who Bears the Working-Capital Burden
Strong Monthly Industry Conditions Do Not Mean Every Segment Has the Same Profit Sensitivity
Over the Next 2–4 Quarters, 8 Indicators Matter More Than Price Targets
5 Disconfirming Factors That Will Determine How Far This Thesis Can Run
Conclusion: AI Is Moving from “Computing More” to “Doing More”
本内容基于公开资料和研报数据整理,不构成任何投资建议,不代表任何个人观点,仅供学习参考,请理性阅读
The next leg of growth in AI hardware will come from more than just additional GPUs. As agents begin performing real work, CPUs, general-purpose servers, memory, and power infrastructure are all entering an expansion cycle, while consumer electronics are paying the price for the same supply shortage.
TL;DR
The most important conclusion in Nomura’s July monthly report is that as AI moves from training to agent execution, the hardware configuration is shifting from approximately 1 CPU core per 8 GPU units toward a CPU-to-GPU ratio closer to 1:1. GPUs are not being replaced; what is being added is CPU orchestration demand arising from planning, branching, tool calls, state management, and waiting for external services.
This wave of CPU demand is not confined to AI servers. Nomura expects general-purpose server sales to increase from US$105 billion in 2025 to US$181 billion in 2026 and further to US$259 billion in 2027; over the same period, AI server sales are expected to rise from US$188 billion to US$335 billion and US$591 billion. The simultaneous rise of both curves indicates that compute systems are shifting from “standalone accelerator expansion” to “full-stack infrastructure expansion.”
Monthly data have already begun validating this chain. In May 2026, global semiconductor shipments measured by the World Semiconductor Trade Statistics organization increased 118.4% year over year; in June, the combined sales of 265 Taiwanese electronics companies increased 59.6% year over year, with TSMC, Hon Hai, and Quanta growing 67.9%, 52.1%, and 102.9%, respectively; the 3-month moving average of Japanese semiconductor equipment sales increased approximately 18% year over year in May.
Memory provides the clearest price signal of diverging supply and demand. Based on the figures in Nomura’s main report, DDR5 DRAM spot prices increased 16.0% month over month at the end of June, while contract prices rose 2.0%, leaving spot prices at an approximately 133.5% premium to contract prices. AI server customers are willing to accept price increases to secure supply, while PC and smartphone customers are more likely to reduce demand in response to higher prices. Global PC shipments declined 4.9% year over year in Q2 2026, the first negative reading in 9 quarters.
Headlines about massive capacity expansion do not imply a near-term surplus. Nomura believes it will take at least 5—10 years for the long-term investments announced by Korean memory manufacturers to translate into meaningful capacity. The key metrics to monitor next are the CPU-to-GPU ratio, general-purpose server orders, the DRAM spot-to-contract price spread, equipment deliveries, and end-product sales. If agent deployment, power infrastructure, or corporate returns fail to keep pace, the current simultaneous rise in volumes and prices could first cool at the valuation level.
The Real New Insight in This Monthly Report Is Not a Larger CPU Market Estimate
Over the past several months, estimates for the server CPU market have already undergone an intensive round of upward revisions. Different institutions’ estimates of the 2030 total addressable market have risen from US$132 billion to US$170 billion and then US$223 billion, with debate focused on how AMD, Intel, the Arm architecture, and Nvidia will divide the incremental revenue. Continuing to focus on the upper bound of these estimates would risk repeating the same argument.
Nomura’s July monthly report advances the architecture thesis into monthly validation: why agents require more CPUs; whether this shift has already flowed through to server sales, foundry revenue, memory prices, equipment orders, and materials procurement; and why the same supply shortage has begun squeezing PCs and smartphones. The research focus has shifted from “how large the market might become” to “whether demand is actually materializing across the supply chain.”
This distinction matters. Total addressable market estimates are a set of long-term assumptions; any change in configuration ratios, average selling prices, or penetration rates can alter the 2030 result by tens of billions of US dollars. Monthly sales and pricing data are noisier but more closely reflect actual orders. If the CPU resurgence is only a narrative without deliveries, general-purpose server revenue, TSMC’s monthly revenue, semiconductor equipment sales, and memory prices should not all strengthen simultaneously. If these data continue moving in the same direction, repeated upward revisions to long-term market estimates will have a firmer basis in reality.
Nomura’s conclusion is also not that “CPUs will replace GPUs.” Training, prefill, and high-throughput inference still require massively parallel computing. Agents merely transform a single model call into a longer sequence of workflows, causing demand for processing, storage, networking, and power beyond GPUs to increase in parallel. Hardware investment is broadening from a single dominant GPU theme into a wider system-level buildout.
From 1:8 to Nearly 1:1: The Change Is in the Workload, Not the Chip Narrative
Training workloads are relatively concentrated. Large volumes of matrix computations can be executed in parallel, with GPUs handling the primary computation and CPUs mainly responsible for system initialization, data preparation, input/output, and scheduling. Nomura uses approximately 1 CPU core per 8 GPU units to characterize the resource structure of the training-dominated era. This ratio summarizes the GPU’s overwhelming position as the center of demand during the training era and should not be treated as a universal engineering standard for all servers.
Agent execution takes a completely different form. After a user provides an objective, the system must first decompose the task, determine which model or tool to invoke, query databases, connect to enterprise applications, check permissions, wait for external interfaces to return results, decide the next step based on those results, and continuously write back intermediate states. Multi-agent systems must also allocate work among different agents, exchange results, and handle failures and retries. Many steps cannot be parallelized indefinitely because each subsequent step must wait for the output of the preceding one, while conditional branches continuously alter the execution path.
This is precisely where CPUs excel: sequential execution, complex control, low-latency access, and operating-system and application scheduling. GPUs still handle large-model inference, while CPUs are gradually becoming the control plane. A single conversation may trigger only one model response, but an agent task may trigger dozens of model, database, and application calls. The user sees one task; underneath, the system is executing an ever-expanding decision tree.
AMD’s market assessment published in May 2026 is directionally consistent with Nomura’s: conventional chatbot architectures typically use 1 CPU to manage 4—8 GPUs, while agents must repeatedly access databases, interfaces, enterprise applications, and memory systems, potentially driving the total server CPU market above US$120 billion by 2030. The two sources use different levels of granularity, so “CPU cores,” “CPU counts,” and “GPU units” cannot be directly interchanged. The common conclusion that can be confirmed is that CPU configuration intensity relative to GPUs is increasing.
A shift toward a 1:1 ratio has 2 additional implications. First, CPU demand comes not only from the head nodes of GPU racks but also from standalone CPU servers responsible for orchestration, retrieval, databases, and small-model workloads. Second, the boundary between general-purpose and AI servers is becoming blurred. A server without a large number of GPUs is already part of AI infrastructure if it handles agent control, vector retrieval, permissions, and enterprise applications.
Therefore, renewed acceleration in general-purpose servers does not mean AI investment is retreating; it may instead indicate that application-layer deployment is beginning. During the training era, the most visible component of capital expenditure was a small number of hyperscale GPU clusters. In the agent era, demand will spread to enterprise data centers, cloud instances, databases, storage, and general-purpose computing. The direct reason CPUs are moving from a supporting role into the orchestration layer is increasing system complexity, not compute density exceeding that of GPUs.
The Simultaneous Rise of Both Server Curves Is the Core of the “Hardware Relay”
Nomura’s chart divides the server market into general-purpose servers and AI servers. In 2024, sales of the 2 categories were approximately US$85 billion and US$107 billion, respectively; in 2025, they increased to US$105 billion and US$188 billion; in 2026, they are forecast to reach US$181 billion and US$335 billion; and in 2027, they are forecast to rise further to US$259 billion and US$591 billion.
AI servers remain the larger revenue pool, while the share of general-purpose servers continues to decline, so rising CPU demand cannot be interpreted as replacing GPU spending. The real change is the expansion of the overall market: general-purpose servers are expected to grow approximately 147% from 2025 to 2027, while AI servers are expected to grow approximately 214%. The GPU theme continues to expand, while CPU-dominated general-purpose computing is reaccelerating from a low base.
This also explains why multiple companies have simultaneously raised their CPU market expectations. Company estimates compiled by Nomura show that Arm believes the relevant market could reach US$100 billion by 2030; AMD raised its 2030 outlook from US$60 billion in Q4 2025 to US$120 billion in Q1 2026; Nvidia expects CPU sales of approximately US$20 billion in 2026; and Qualcomm estimates that the relevant market could reach US$200 billion by FY2029. The definitions differ: some figures refer to server CPUs, while others include broader processor or data-center opportunities. They therefore cannot be ranked directly, but corporate actions are more informative than a single long-term forecast.
Product roadmaps are also becoming denser. Nvidia Vera, Google Axion, Microsoft Cobalt, and Amazon Graviton are all advancing along the Arm architecture path, while Arm and Qualcomm also plan to introduce new data-center CPUs. The Arm architecture is attractive for its power consumption, cooling characteristics, and customization flexibility, while x86 benefits from mature software, compatibility, reliability, and established enterprise systems. The agent era will not automatically allow one architecture to dominate the entire market; instead, it will segment the market into multiple layers, including high-performance general-purpose computing, cloud-provider customization, GPU coordination, and cost-sensitive deployments.
The sequence of the hardware relay therefore resembles “GPU expansion—CPU reinforcement—storage and power follow-through,” rather than “GPU peak—CPU replacement.” If general-purpose server sales continue growing while AI servers also maintain high growth, the 1:1 ratio represents incremental system demand. If general-purpose server growth is driven by customers purchasing in advance and AI server growth subsequently slows, the purported relay may merely reflect inventory and budgets shifting among different categories.
Monthly Data Are Turning the Architecture Narrative into Revenue
The first set of corroborating evidence comes from global semiconductor shipments. Preliminary data for May 2026 cited by Nomura show that global semiconductor shipment value rose 118.4% year over year, with the 3-month moving average up 104.5%. The Americas, Europe, Japan, and Asia-Pacific grew 133.5%, 60.3%, 22.1%, and 103.5%, respectively. Such high growth is not entirely volume-driven: memory price increases, a higher share of advanced logic, and volume ramp-ups of high-value AI chips all lift shipment value.


