目录
I. The Inflection Point in AI-Lab Compute and Revenue
II. Two Labs Are Becoming the Centers of Global Compute
III. Semiconductor Supply Chains and the Limits of Capacity Expansion
IV. The 100 GW Target for 2028
V. Model-Release Constraints and Revenue per MW
VI. Who Ultimately Captures the Value of AI?
VII. The Bullwhip Effect in Supply-Chain Pricing
VIII. Training Will Consume More Compute
9. The Global Compute Landscape Beyond 2028
10. Export Controls, the Financial System, and the Pace of Catch-Up
11. The Compute Mix Across Research, Development, and Inference
12. $1 Trillion in Capex and Front-Loaded Construction
13. Credit Markets Will Become the Constraint on Expansion
14. How AI Investment Could Raise Economy-Wide Interest Rates
15. How Much Debt Will the $11 Trillion Buildout Require?
16. 8% Interest Rates and a Second Volcker Shock
17. The Opportunity Cost of Capital in a High-Growth Economy
18. AGI Is Constrained by More Than Research Capability
19. The AI Workforce Is Expanding Tenfold Every Year
20. Why Economies of Scale Naturally Drive Concentration
21. Can External Value Capture Provide a Buffer?
本内容基于公开资料和研报数据整理,不构成任何投资建议,不代表任何个人观点,仅供学习参考,请理性阅读
I. The Inflection Point in AI-Lab Compute and Revenue
Host: Today we are once again joined by SemiAnalysis founder Dylan Patel. What we call our “family Thanksgiving dinner” is really just our annual podcast recording. Of course, we are not actually related—but don’t tell anyone, or the legend will be ruined.
The direction of the global economy increasingly depends on AI labs’ economics, the evolution of the compute market, and everything surrounding them. I want to understand where this extraordinary future is heading over the next several years. But let’s begin with the present. Walk us through the current compute and revenue picture for AI labs, and what you expect over the next year or two.
Dylan: Looking back at last year, AI infrastructure accounted for a large share of US GDP growth, even by year-end.
Of all the compute capacity added this year, roughly one-third will ultimately be used by OpenAI and Anthropic. Other companies may build the infrastructure and lease it to them, but these two labs remain the end customers.
Compute volumes will continue to expand rapidly. Capital expenditure this year is slightly above $1 trillion and will exceed $2 trillion by 2028, with the labs accounting for an increasingly large share.
This will ultimately create a remarkable situation: the labs’ annual spending will rise from tens of billions of dollars to hundreds of billions. Based on some of the contracts they have already begun signing with partners, projected annual spending reaches trillions of dollars by the end of this decade.
That requires a major change in their economics. Until recently, these were essentially companies that continuously lost money. Anthropic became profitable in the second quarter. OpenAI is also expected to reach profitability at some point in the third quarter as products such as Codex and 5.6 continue to grow.
A year ago, their losses were funded entirely by venture capital. That was still the case at the beginning of this year. They have now crossed the inflection point and begun generating real profits.
That does not mean they have stopped raising capital. New funding is still flowing in to accelerate growth further, but an increasing share of operating expenditure will be funded by their own revenue rather than relying entirely on external capital injections.
Their margins have improved dramatically over the past year and a half. The underlying cost of compute capacity is typically around $10 million–$15 million per megawatt.
The most important change is that OpenAI previously generated negative gross margins when serving GPT4 inference on Nvidia Hopper GPUs. Today, whether OpenAI is serving GPT56 or Anthropic is serving Opus 5 or Fable 5, revenue per megawatt far exceeds the $10 million–$15 million cost of adding each megawatt of compute.
Anthropic has generated as much as $50 million in revenue per megawatt. In other words, every $10 it spends building inference capacity can produce $50 of revenue, allowing it to invest all the resulting incremental profit in training.
II. Two Labs Are Becoming the Centers of Global Compute
Host: I am very interested in how you view the concentration of compute among a small number of labs, and how the global balance of compute between these labs and everyone else will change.
If roughly one-third of new compute currently goes to these labs, when will they receive more than half of all new global compute? And when will they control the vast majority of global compute?
Dylan: At the beginning of this year, OpenAI had 2 GW of compute and Anthropic had less than 2 GW. By year-end, each will exceed 5 GW, implying that their aggregate compute capacity will have increased three- to fourfold.
They account for roughly 30% of the compute added this year. Next year, based on contracts that have already been formally signed, the shift will be even more pronounced. Anthropic and OpenAI could absorb 40%–50% of all new compute capacity.
This concentration does not appear likely to slow or stop; if anything, it will continue to accelerate. The parties building capacity for them will change. SpaceX, for example, will become a major new entrant next year and build substantial compute capacity.
SpaceX will probably lease a significant portion of that capacity to Anthropic and OpenAI because those two companies have the highest marginal willingness to pay.
At the same time, OpenAI and Anthropic have begun building their own compute capacity. OpenAI is developing proprietary chips, while Anthropic is purchasing TPUs from Google and deploying them in partnership with Fluidstack.
So if the question is when half of all new global compute will go exclusively to OpenAI and Anthropic, the answer is by the end of next year. At that point, the two companies will account for half of all incremental capacity.
Because compute capacity is growing so rapidly, this new capacity will soon represent most of the global installed base.
Host: In other words, within perhaps just a year and a half to two years, most global compute will either be owned by these two labs or at least dedicated to meeting their needs.
The current trend seems to be that global gigawatt-scale compute doubles every year, while frontier-lab compute triples annually.
If that trend continues, the figure rises from 2 GW at the beginning of this year to nearly 6 GW by year-end. Multiplying by three gives 18 GW by the end of 2027 and 54 GW by the end of 2028.
At that point, global supply constraints would clearly prevent them from continuing to triple capacity every year. How do you view global compute supply over the next several years?
Dylan: If 30 GW of capacity is added this year, 50 GW next year, and roughly 70 GW the following year, an important dynamic emerges: every watt deployed this year is substantially more efficient than capacity deployed two years ago.
A large share of the world’s existing compute capacity was installed this year. Even if installed power capacity has not doubled, the equipment being deployed is GB300, TPUV7, and TRAINIUM3. These are far more energy-efficient than the previous generation, delivering three to five times the performance per watt.
That creates a massive step-up in effective performance.
Assume Anthropic and OpenAI absorb 45% of next year’s new capacity. By December 2027, they would account for roughly half of global incremental compute. Yet the actual performance of that new capacity would exceed that of all previously deployed capacity, requiring an additional performance multiplier.
By the end of 2028, if this trend continues—and I currently see nothing that would stop it—the two companies will directly control most of the world’s genuinely usable floating-point compute.
III. Semiconductor Supply Chains and the Limits of Capacity Expansion
Host: What I do not understand is this: if the value of compute is rising so quickly, why do you believe only around 80 GW can be added in 2028?
Dylan: Incidentally, that is already the upper limit—an extremely optimistic scenario.
Host: Then let’s work through the supply chain.
When I interviewed you a few months ago, you said that producing 1 GW of Vera Rubin compute would require roughly 55,000 N3 wafers, 6,000 N5 wafers, and 170,000 DRAM wafers. Those figures may have changed since then.
As an aside, I joked at the time that your pronunciation of “wafer” sounded distinctly Indian.
Dylan: When we first moved to the US, I genuinely struggled to distinguish between “v” and “w.” I was also a vegetarian at the time.
Host: You have told me that story before. When you attended elementary school in North Dakota, you pronounced “veggie” as “wedgie.”
But returning to the point, those are the wafer volumes required to produce 1 GW of compute. I had a large model run your fab-equipment model to calculate the equipment investment needed to produce 1 GW annually. The result was $3 billion–$4 billion. Including cleanrooms, the building shell, and other fab facilities, the total might be around $6 billion of fab capital expenditure.
In other words, $6 billion of fab capital expenditure can support annual production of 1 GW of compute, while 1 GW of compute can currently generate $100 billion in revenue.
That $6 billion investment does not produce only one year’s output. It can manufacture 1 GW every year, with each gigawatt generating $100 billion in annual revenue.
Over five years, the 1 GW produced in the first year can generate five years of profit, the 1 GW produced in the second year can generate four years of profit, and so on. Ultimately, $6 billion of fab-level capital expenditure can support more than $1 trillion of end-market AI revenue.
Dylan: Yes, but there are substantial operating expenses and other capital expenditures in between, including data centers, power, and installation. OpenAI’s R&D; also has to be funded. Many participants throughout the chain need to earn revenue.
Host: Even if half the revenue goes to those intermediate stages, the gap between fab capital expenditure and end-market revenue is still roughly 100-fold. The actual gap may be wider; we are using very conservative assumptions.
That is capitalism. If there is an enormous opportunity to turn $1 into $100, won’t the supply chain find a way to produce more mirrors?
Dylan: It will, but manufacturing those mirrors takes time.
OpenAI and Anthropic will soon say: “We could be generating $1 trillion today, but we are constrained by the mirrors inside ASML equipment. If we invest $100 billion, how can we manufacture more mirrors?”
Host: I find it difficult to imagine that this supply constraint will ultimately remain unresolved.
Some interesting arbitrage has already emerged. People are buying gas turbines in advance and then trying to resell them because gas turbines have become a bottleneck for data centers, making them worth far more than their original price.
If someone has $400 million and can persuade ASML to sell them an EUV lithography system, they should simply buy it, wait, and then resell it for more than $1 billion.
Dylan: Capitalism will, of course, ultimately drive capacity expansion. But this is a very long bullwhip: demand signals take considerable time to propagate to the far end of the supply chain.
The supply chain does not respond immediately. Ask Carl Zeiss today, and it will say: “Yes, by the end of this decade, we need to produce enough mirrors for 100 EUV systems annually.”
When we recorded our program earlier this year, the company did not even believe it needed to reach that output level. It is only now beginning to accept a target of 100 systems per year.
Given the current economic returns, actual demand should be even higher. The problem is simply that capacity expansion takes too long.
Host: Suppose every company across the supply chain were acquired by private equity firms that firmly believed AGI was imminent and decided to maximize output. What physical constraints would prevent further production expansion?
I ask because the labs’ revenue—and the cash flow of the entire AI industry—will soon become extraordinarily large. Accelerators themselves will also generate enormous cash flow. At that point, extreme supply-chain expansion could be funded entirely from internally generated cash.
Dylan: I agree in principle. Physical constraints would, of course, remain. At the supply chain’s current pace of expansion, roughly 100 ASML systems annually by 2030 remains a reasonable forecast.
But the outcome would change if you went directly to Carl Zeiss and said: “Here is $10 billion. Expand capacity immediately.” The problem is that you would have to do the same with every company across the supply chain.
Host: You do not think that will happen next year?
Dylan: I do not think it will happen this year, next year, or the year after, because the entire world remains capital-constrained.
Host: But if the frontier labs collectively generate $1 trillion in revenue next year, could they not allocate $10 billion—or even hundreds of billions of dollars—to expanding the supply chain?
Dylan: I do not think they will.
The labs will generate hundreds of billions of dollars in revenue next year, but total capital expenditure could reach $2 trillion, leaving an enormous funding gap.
The wafer-fabrication equipment supply chain may require roughly $200 billion. The data-center supply chain will require more, the accelerator supply chain even more, and the energy supply chain will also need massive investment. Taken together, capital expenditure will far exceed $2 trillion.
The labs’ cash flow therefore remains insufficient to fund all the required expansion.
Host: Of course, they may never rely entirely on cash flow, because the rational strategy is to keep capital expenditure above current returns and continue reinvesting.
IV. The 100 GW Target for 2028
Host: What I am really trying to understand is this: if current trends persist, each lab will have more than 50 GW of compute capacity by the end of 2028, bringing the two to a combined 100 GW.
And as you noted, continued hardware improvements mean that by 2028, each gigawatt will deliver several times today’s throughput and performance. This is not just about higher floating-point performance per watt; the hardware will also process AI workloads more effectively.
So how much total compute capacity will the world have in 2028?
Dylan: Reaching a combined 100 GW could be extremely difficult, because by 2028 the two labs would need to absorb 70%–80% of all incremental global compute capacity.
I am not sure how the market would respond. How high would compute prices have to rise before they could actually secure 70%–80% of new supply? Would Google, Meta, and Amazon be willing to sell that much compute?
There is also a classification issue. If Amazon serves Anthropic models through Bedrock, we still count that as Anthropic compute because the ultimate revenue is recognized by Anthropic, even though the arrangement may include revenue sharing, credit rebates, and other mechanisms.
In any event, if the two companies reach a combined 100 GW in 2028, they will necessarily have caused enormous market disruption.
Today, it is genuinely not that difficult for anyone to make money using compute that costs $10 million–$15 million per MW.
You can buy a GB300 rack, download the Kimi weights, then download vLLM or SGLang and deploy the system. Codex and Fable can even help you do this. It is not effortless, but it is hardly rocket science. Put the service on OpenRouter, and you can begin generating revenue above the cost of the compute.
That is precisely why prices for compute costing $10 million–$15 million per MW have already begun to rise.
To believe that the two labs can control 100 GW in 2028, you must also believe they can outbid other buyers, because anyone can make money at $10 million–$15 million per MW.
The question is whether compute prices rise to $25 million per MW—or even $40 million per MW.
Host: It is already clear that the labs generate far more revenue per MW than other participants. If they maintain their current lead, that advantage could theoretically persist.
If some form of recursive self-improvement relatively strengthens the labs’ R&D; capabilities—or if they have unreleased internal models that help develop the next generation—their advantage could widen further.
Are we already seeing this? If SpaceX or other slightly lagging participants cannot monetize compute internally as efficiently as the labs, they will sell it to the highest bidder. The labs will therefore keep bidding for a larger share of the market.
Dylan: That is also my base case. They will continue absorbing an ever-larger share of compute.
But ultimately, they cannot achieve this at current or near-current prices. To absorb 70% of global compute and reach 100 GW in 2028, they would have to start paying $25 million, $30 million, or even $50 million per MW. That is an extremely aggressive target.
V. Model-Release Constraints and Revenue per MW
Dylan: Another major challenge is that AI-lab progress has already slowed materially.
The regulatory measures they advocate are actually slowing the labs more than they are affecting China’s open-source language-model market.
OpenAI did not release Astra and paused training for two weeks. Anthropic likewise did not release the model called Model 2 in its safety evaluations, which is widely believed to be the next version of Mythos.
They have clearly withheld their strongest models. Under those conditions, revenue per MW will stagnate or even decline again as competing models regain competitiveness.
This does not mean the labs have fallen behind; they simply have not released their best work.
If regulatory constraints prevent them from releasing their strongest models, revenue per MW will not rise at its previous pace. Their ability to outbid others for incremental compute will weaken, potentially preventing them from reaching 100 GW.
But in a world without safety constraints, I do believe the earlier scenario would play out. They could generate $100 million or more in revenue per MW and afford to pay $50 million per MW.
At that point, other compute owners would have no reason to deploy their capacity elsewhere. They would simply say: “Dario, take all of my compute.”
In reality, however, several forces could slow this process.
Host: One useful thought experiment is: what happens if AI models truly reach the level of fully autonomous software engineers?
They are not there yet. In my view, they remain far from fully automating the work of an entire white-collar employee.
But white-collar employees typically earn $100,000 or more per year. If 1 GW of compute could sustain an AI workforce equivalent to roughly 1 million white-collar employees, that would represent $100 billion of value.
Dylan: At $100,000 per person across 1 million people, that is indeed $100 billion. If anything, the figure is surprisingly low.
With full AGI, the value created per GW could reach hundreds of billions of dollars.
However, we continue to observe another phenomenon: OpenAI and Anthropic do not capture most of the value their models create. Fortunately, most of that value has so far accrued to users.
For example, Jane Street has an exclusive contract with OpenAI for GPT-5.6 Ultrafast mode and is also one of Anthropic’s largest customers. The value Jane Street extracts from the tokens it buys far exceeds Anthropic’s own profit because Jane Street can use the models to make money in the market.
Consider Meta, which was reportedly responsible for 10% of Anthropic’s business at one point. Meta uses these models to optimize its advertising algorithms, improving efficiency and user engagement. Even a 5% increase in engagement would generate far more money for Meta than for Anthropic.
That is the premise on which the entire system works.
Host: Of course, adding 1 million software engineers would also reduce the market cost of software engineers.
What puzzles me is whether the market ultimately reaches equilibrium. If it does, will compute prices converge toward the revenue Anthropic and OpenAI can generate per MW, leaving the labs only a modest markup?
The current situation is unusual: there is a fourfold or greater gap between the price at which compute is sold and the revenue Anthropic earns from using it.
If revenue per GW continues increasing and Anthropic’s monetization of 1 GW rises another two or three times, it would be anomalous for that gap to keep widening. With nothing more than a set of model weights, Anthropic can turn something that costs $10 into $100.
VI. Who Ultimately Captures the Value of AI?
Dylan: The question of where AI’s value ultimately accrues has always been fascinating.
AI is creating enormous value. First are end users, who, I think we agree, receive more value than any other layer—that is why they are willing to pay so much for the models.
Next is the application layer, which so far has created or captured very little value.
Below that is the model layer. A year ago, its gross margins were negative. It is now generating substantial positive gross profit and appears to be moving toward $100 million of revenue per MW—turning $10 million–$15 million into $100 million.
But a year ago, virtually all gross profit was captured by the hardware supply chain, while everyone else lost money. OpenAI, Anthropic, and many startups continued deploying venture capital, while hyperscalers built infrastructure without knowing whether they would ultimately earn a return.
At the time, the model layer was effectively creating negative value because token revenue was below infrastructure costs. All the value was captured by chipmakers and foundries.

