目录
I. The Bottom Line: The Same Model Business Can Produce Two Very Different Income Statements
II. The First Hurdle: Put Gross versus Net Reporting and ARR Definitions on a Common Basis
III. Three Product Models: Subscriptions, Direct APIs, and Indirect APIs Monetize Differently
IV. The 2026 Margin Inflection: Lower Costs Matter More Than Higher Model Prices
5. Why Cloud Providers Remain Major Winners: They Monetize Compute, Distribution, and Delivery Certainty
6. Training Costs Change the Conclusion: High Inference Margins Do Not Guarantee High Free Cash Flow
7. From $7 Billion to $690 Billion: The Forecast’s Structure Matters Most
8. The 2028 Inflection Point: Labs Will Move More Compute Beyond General-Purpose Public Clouds
IX. Who Has the Edge: Focus on Contracts, Distribution, and Utilization—not a Single Model Leaderboard
X. The Strongest Counterevidence and Key Metrics to Track: Five Sets of Numbers Will Determine Whether the Framework Holds
本内容基于公开资料和研报数据整理,不构成任何投资建议,不代表任何个人观点,仅供学习参考,请理性阅读
The defining shift in 2026 is the sharp decline in unit inference costs: AI labs are rebuilding gross margins, while cloud providers continue to capture high-margin compute revenue. Winners will be determined by their business mix, accounting treatment, and the infrastructure transition around 2028.
I. The Bottom Line: The Same Model Business Can Produce Two Very Different Income Statements
AI lab unit economics have entered a new phase. Barclays estimates that inference compute accounted for $61—79 of every $100 of revenue in 2025; by 2026, that figure had fallen to approximately $30. Paid-inference margins rose from 14%—18% to 48%—65%. After deducting the training costs of production models, the two hypothetical labs achieved blended gross margins of 38% and 55%, respectively.
These figures show that improvements in model capability are beginning to translate into commercial efficiency. Quantization, speculative decoding, batching, caching, routing optimization, and higher utilization among enterprise customers have collectively reduced the compute cost of each effective task. Model prices have not declined nearly as sharply, allowing the cost savings to accrue first to lab gross margins. Growth in enterprise agent workflows has also improved the customer mix: high-utilization, billable workloads with defined usage caps now represent a larger share, reducing the drag from free users and inefficient interactions.
The same underlying improvement does not automatically produce comparable financial statements. Revenue recognition, channel costs, and partner revenue shares are treated differently across subscriptions, direct API sales, and indirect API sales through cloud marketplaces. One lab may recognize cloud-marketplace transactions on a gross basis, while another records only the amount net of channel fees. One may classify strategic-partner revenue sharing as cost of revenue, while another may not recognize revenue hosted by its cloud partner at all. Reported revenue growth, gross margin, and R&D; expense ratios therefore cease to be directly comparable.
Investors should compare three layers of economics: what customers pay for model services; what the lab retains after inference and channel costs; and how much sustainable profit and cash flow remains after production-model and R&D; training expenses. GAAP financial statements remain important, but the underlying contracts must first be mapped onto the same value chain.
II. The First Hurdle: Put Gross versus Net Reporting and ARR Definitions on a Common Basis
Accounting differences can make two operationally similar companies appear to be at entirely different stages of development. Barclays illustrates the issue with Lab A and Lab B. The two labs serve similar customer demand and incur comparable underlying inference costs, yet differences in revenue recognition and strategic-partnership terms produce wide gaps in reported revenue, costs, and gross margins.
Indirect API sales are the most prone to misinterpretation. When customers purchase model services through AWS, Azure, or Google Cloud, the cloud provider handles billing, sales, and customer-relationship management. If the lab is the principal in the transaction, it may recognize $100 of gross revenue and then pay $20 in channel or cloud fees. If the cloud provider is the principal, the lab may recognize only $80 of net revenue. Customer spending is identical, but the lab’s reported revenue differs by $20.
Strategic partnerships introduce another layer of divergence. Barclays assumes that Lab B pays a strategic partner a 20% revenue share, while Lab A has no such provision. Lab B’s underlying product economics may be no worse, but the revenue share reported on its income statement reduces both paid-inference margin and blended gross margin. If partner-hosted revenue is excluded entirely from Lab B’s financial statements, the market may also underestimate its true end demand and model usage.
ARR also requires disaggregation. Annualized subscription contract value, current-period API consumption, committed cloud-marketplace spending, and strategic-partner-hosted revenue may all be annualized differently. Page 2 of the report even uses customer spending of $101 and $102 to show that the two labs define $100 of reported revenue differently. Comparing disclosed ARR directly can therefore mistake revenue-recognition differences for shifts in market share.
A more reliable framework tracks four metrics: end-customer spending, paid Tokens or effective task volume, contribution profit after inference costs, and cash allocation between the lab and its cloud partners. Once these four measures are aligned, the apparent gap created by gross versus net revenue reporting narrows substantially. As AI labs begin providing more comprehensive GAAP disclosures, the market may initially enter a phase in which more data paradoxically reduces comparability.
III. Three Product Models: Subscriptions, Direct APIs, and Indirect APIs Monetize Differently
Subscriptions sell reliable access and a consistent user experience. Users pay a fixed monthly or annual fee, while the lab absorbs actual Token consumption. Tools such as Claude Code and Codex control costs through usage allowances, concurrency limits, model routing, and enterprise contract caps. Barclays estimates that mature subscription products generate inference margins of approximately 70%. This is already a high level, but it remains sensitive to power users, free trials, quota resets, and customer-acquisition subsidies.
The subscription model offers predictable revenue, strong product stickiness, and a natural entry point into customer workflows. Its risk is asymmetric usage: a small group of intensive users can consume substantial compute, while fixed pricing cannot immediately reflect the associated cost. Labs must continually adjust pricing tiers, rate limits, and model segmentation. Enterprise contracts with explicit usage caps and overage charges should deliver more stable margins than unlimited consumer plans.
Direct APIs monetize measurable model calls. Developers such as Cursor and Figma purchase Tokens directly from the lab, with pricing tied to usage. Barclays says paid-inference margins for some direct APIs exceeded 80% in Q2 2026, although the report uses a more conservative 70% assumption. Direct APIs translate inference-efficiency gains into profit more quickly because unit pricing is transparent and cost savings do not need to be passed on to customers immediately.
These high margins still have limits. API pricing may fall as competition among frontier models intensifies; routing to open-source and smaller models will reduce the share of calls handled by premium models; and larger customers will demand volume discounts. Current margins include premiums derived from scarce compute, superior capabilities, and switching costs. Long-run normalized margins will probably be below peak levels.
Indirect APIs monetize channel reach. Cloud providers integrate models into their existing procurement, compliance, billing, and sales systems, helping labs reach large enterprises. Labs exchange channel fees or revenue sharing for lower customer-acquisition costs and faster scaling. The economics depend on bargaining power: when models are scarce, labs can retain a larger share; as models become more substitutable, cloud providers can more easily compress that share and control the customer relationship.
Business mix therefore determines the shape of the income statement. Barclays assumes that Lab A’s revenue mix is 30% subscriptions, 45% direct API (API45), and 25% indirect API (API25); for Lab B, the mix is 80% subscriptions, 10% direct API (API10), and 10% indirect API (API10). Both incur approximately $30 of inference costs, but Lab A pays higher channel fees without a strategic-partner revenue share, generating $65 of paid-inference profit. Lab B pays a $20 partner revenue share and generates only $48. Despite similar product efficiency, contract structures create a 17-percentage-point margin gap.
IV. The 2026 Margin Inflection: Lower Costs Matter More Than Higher Model Prices
The shift from 2025 to 2026 was driven almost entirely by lower inference costs. Lab A’s inference compute cost per $100 of revenue fell from $79.1 to $30, while Lab B’s declined from $61.1 to $30. Their paid-inference margins rose from 14% and 18% to 65% and 48%, respectively. Even after deducting $10 of production-model training costs, blended gross margins increased from 4% and 8% to 55% and 38%.
Cost improvement has followed three paths. First, greater model-level efficiency means that the same task requires fewer Tokens or can achieve similar quality at lower precision. Second, serving-layer gains—including batching, caching, parallel scheduling, speculative decoding, and higher chip utilization—reduce compute consumption per Token. Third, the demand mix has improved: enterprise APIs and agent tasks carry higher paid-usage rates, reducing the drag from free interactions on the revenue denominator.
Agent workflows are especially important. Ordinary chats often involve only a few rounds of interaction, whereas agents repeatedly call models, tools, and external systems. Each task consumes more Tokens, and customers are more willing to pay for measurable productivity gains. Higher call volumes increase compute demand, but contribution profit can continue expanding as long as efficiency improves faster than unit prices decline.
The margin inflection should not be extrapolated mechanically. The report also cautions that current levels may exceed the long-run steady state. Competition at the model frontier will push prices lower, expanding compute supply will erode scarcity premiums, and larger enterprise customers will negotiate discounts. Efficiency gains will ultimately be redistributed among labs, cloud providers, and customers. In the near term, they accrue first to lab income statements; over time, some will likely be passed on to customers.
Testing this thesis requires tracking whether unit costs are falling faster than unit prices. Public disclosures may provide five relevant signals: price per million Tokens, Token consumption for the same benchmark task, cache-hit rates or batching ratios, inference compute cost as a percentage of revenue, and the margin differential between APIs and subscriptions. If call volumes continue to rise but the inference-cost ratio stops declining, efficiency gains may be approaching a plateau—or price competition may be consuming the cost savings.
5. Why Cloud Providers Remain Major Winners: They Monetize Compute, Distribution, and Delivery Certainty
Improving gross margins at AI labs have not diminished the value of cloud providers, at least at this stage. Barclays estimates that roughly $35—$40 of every $100 in lab revenue will flow to hyperscalers in 2026, generating $10—$20 in operating profit for the cloud providers and implying an operating margin of approximately 35%—45%. Partnership terms may alter the accounting allocation, but the underlying Token economics could be similar.
Cloud providers capture returns in three ways. First, IaaS revenue, as labs directly purchase GPUs, networking, and storage. Second, distribution revenue, as cloud marketplaces resell model APIs and charge fees. Third, strategic-partnership revenue sharing, through which cloud providers exchange capital, compute commitments, and go-to-market access for a share of lab revenue. All three may coexist, making it difficult to isolate AI’s contribution from headline cloud growth alone.
What cloud providers really sell is delivery certainty. Training and serving frontier models require large-scale clusters, reliable networks, sufficient power and cooling, and security and compliance frameworks already accepted by global enterprise customers. Labs can optimize their models, but they cannot replicate this infrastructure quickly. Cloud providers can also share data centers, power procurement, and sales teams across AI demand and traditional cloud workloads, lowering marginal customer-acquisition and operating costs.
This advantage still depends on utilization. GPU procurement and data-center construction incur capital expenditure upfront, while revenue is realized gradually. If lab demand falls short of commitments, cloud providers may have contractual protection; if end-customer demand falls short of lab expectations, the entire value chain faces low utilization and depreciation pressure. A high operating margin does not automatically translate into a high return on capital, particularly when investment in new data centers, power, and networking grows faster.
Competitively, cloud providers with frontier-model partners, ample power, and mature enterprise distribution are better positioned to capture AI revenue. The report assumes that by 2028, AWS and Google Cloud will each account for approximately 30% of hyperscaler AI revenue, with other clouds representing the remaining 40%. Azure is not shown separately and is included in other clouds. Readers should therefore not interpret “other clouds” as a collection of small providers; it may include platforms as large as Microsoft.
6. Training Costs Change the Conclusion: High Inference Margins Do Not Guarantee High Free Cash Flow
Inference gross margin answers only how much profit remains from serving one additional paying customer. It does not capture the full cost of maintaining model competitiveness. Training next-generation models, acquiring data, hiring researchers, absorbing failed experiments, and optimizing models before deployment all consume substantial cash. Barclays assumes that approximately 10% of training costs relate to the final production model and are recognized in cost of sales, while the remaining 90% is recorded as R&D; expense.
This allocation directly affects gross margin. The more production-training costs recognized in cost of sales, the lower the gross margin; allocating more training costs to R&D; makes gross margin appear higher, although operating profit is still affected. Labs may classify model training, post-training, fine-tuning, and production deployment differently. Gross-margin comparisons must therefore be assessed alongside R&D; expense, capitalization policies, and cash expenditure.
Barclays’ industry model projects training costs rising from $7 billion in 2024 to $210 billion in 2028. As a share of AI lab revenue, training costs decline from 96% to 30%, indicating that their fixed-cost characteristics become more evident at scale. Absolute spending, however, continues to grow rapidly. Even if inference generates high contribution margins, the continuing race to develop frontier models could absorb most of the resulting cash.
Inference and training costs also follow different cycles. Inference costs rise with usage, while unit costs can decline through higher utilization and Serving optimization. Training costs are more step-like: each new model generation requires new clusters, data, and R&D; investment. If gains in model capability begin to slow, the marginal return on training investment will decline. Genuine operating leverage emerges only if application revenue grows fast enough for training costs to fall as a share of revenue.
High-quality earnings at AI labs therefore require three conditions: stable margins on paid inference; training costs growing more slowly than revenue; and R&D; investment that produces durable model differentiation or product stickiness. Looking only at API gross margins risks overstating cash-generation capacity, while focusing only on R&D; losses may understate the contribution-profit foundation already established by inference.
7. From $7 Billion to $690 Billion: The Forecast’s Structure Matters Most
Barclays expects total AI lab revenue to rise from $7 billion in 2024 to $26 billion in 2025, $137 billion in 2026, $376 billion in 2027, and $690 billion in 2028. The corresponding growth rates are 276%, 418%, 175%, and 84%. Year-end AI ARR is projected to increase from $44 billion in 2025 to $782 billion in 2028.
The absolute forecasts are highly aggressive, but the structural changes are more analytically useful. In 2026, accelerating revenue growth coincides with improving unit economics. Revenue grows 418%, while the inference cost ratio falls from 66% in 2025 to 42% and the training cost ratio declines from 70% to 48%. The combination of scale and efficiency explains the rapid recovery in lab gross margins.
Revenue continues to grow rapidly in 2027—2028, although growth begins to decelerate. The inference cost ratio remains at 42%, implying that subsequent efficiency gains are largely passed through in pricing or consumed by more complex workloads. The training cost ratio continues to decline from 35% to 30%, providing the principal source of operating leverage. If the actual inference cost ratio is below 42%, lab profits could exceed the forecast; if price competition intensifies, both revenue and gross margin will come under pressure.
Hyperscaler AI revenue is projected to rise from $11 billion in 2024 to $36 billion in 2025, $124 billion in 2026, $289 billion in 2027, and $502 billion in 2028. As a share of AI lab revenue, it declines from 153% to 73%. The ratio exceeds 100% in the early years because of training compute, prepaid commitments, and revenue not yet recognized by the labs; it gradually falls as end-market lab revenue expands.
This ratio trajectory matters more than any single revenue figure. It describes the AI value chain’s transition from “build compute first, generate application revenue later” toward a model in which application revenue covers infrastructure investment. The declining ratio does not mean cloud AI revenue is shrinking—it continues to grow. Rather, cloud spending per $1 of lab revenue gradually declines as lab efficiency and bargaining power improve.
8. The 2028 Inflection Point: Labs Will Move More Compute Beyond General-Purpose Public Clouds
Barclays views 2028 as a critical inflection point in infrastructure architecture. The report defines “floor-backed AI infrastructure” as compute projects built by labs, capital partners, or specialized infrastructure providers, with demand underwritten through long-term purchases, minimum-usage commitments, or financing arrangements. These projects can reduce labs’ dependence on general-purpose public clouds.
Scale is the key driver. When lab usage is limited, public clouds offer clear advantages in flexibility, deployment speed, and global reach. Once spending reaches billions of dollars, fixed procurement, dedicated clusters, and long-term power contracts may deliver lower unit costs. Labs also want greater control over chip configurations, network architecture, and scheduling software while reducing cloud distribution fees and strategic revenue sharing.
The migration will not happen all at once. Training, peak inference demand, regional compliance, and enterprise cloud marketplaces will continue to require public clouds. A hybrid structure is more likely: large, stable baseload workloads will move to dedicated facilities, while highly variable, geographically dispersed workloads—or those requiring enterprise distribution—will remain on hyperscale cloud platforms. Cloud providers’ share may decline even as their absolute revenue continues to grow.
Dedicated infrastructure also introduces new risks. Labs must assume longer asset cycles, financing costs, and fluctuations in utilization. If models or products evolve faster than data centers depreciate, hardware may become obsolete prematurely. If end demand falls short of expectations, minimum-purchase commitments will shift from revenue protection for cloud providers to fixed obligations for labs.
Cloud providers can respond by lowering unit prices, offering custom chips, strengthening model-marketplace distribution, integrating enterprise data and security capabilities, and participating in the financing of dedicated projects. Competition will shift from “who owns the most GPUs” to “who can consistently deliver usable compute at the lowest total cost while retaining the customer relationship.”
IX. Who Has the Edge: Focus on Contracts, Distribution, and Utilization—not a Single Model Leaderboard
AI labs’ competitive advantage will ultimately rest on three capabilities. First, inference efficiency: delivering equivalent performance with less compute. Second, product mix: enterprise APIs and workflow products are generally more likely to generate sustainable profits than subsidized consumer subscriptions. Third, distribution bargaining power: whether labs can retain customer ownership, reduce partner revenue sharing, and allocate workloads across multiple clouds.
Hyperscalers likewise have three critical capabilities. First, access to power, chips, and networking, which determines deliverable capacity. Second, utilization management, which determines whether heavy capital expenditure can generate adequate returns. Third, enterprise distribution: cloud marketplaces, identity, security, data, and compliance capabilities can reduce friction in model procurement.
Model leaderboards still influence near-term market share, but commercial outcomes depend more on consistent delivery. A leading model may struggle to sustain a revenue advantage if it is too costly, has unstable latency, or lacks enterprise controls. A slightly less capable model with better cost, reliability, and distribution may achieve stronger unit economics. Open models will accelerate capability diffusion, forcing labs to shift value creation toward tools, data, workflows, and customer relationships.
As accounting disclosures mature, the market may reassess both groups. Labs must demonstrate that rapid growth can translate into contribution profit sufficient to cover training and R&D; investment. Cloud providers must show that AI revenue can cover depreciation, financing, and power costs. Both may report substantial AI revenue, but returns on capital are much harder to disguise.
This report provides neither stock-specific price targets nor valuation multiples, so it cannot support precise valuation conclusions. Instead, it offers a profit-pool framework: labs capture model and product profits, while cloud providers capture compute, distribution, and strategic-partnership economics. As dedicated infrastructure matures, these profit pools will be redistributed. Any valuation judgment requires additional company-level data on revenue, capital expenditure, depreciation, contract duration, and cash flow.
X. The Strongest Counterevidence and Key Metrics to Track: Five Sets of Numbers Will Determine Whether the Framework Holds
The first is the inference cost ratio. If AI-lab revenue continues to grow while inference costs as a share of revenue begin rising again, model complexity, price competition, or inefficient usage is consuming the efficiency gains. The report assumes that the inference cost ratio remains stable at 42% from 2026—2028, making this the framework’s most directly falsifiable assumption.
The second is the training cost ratio. The report expects it to decline from 48% in 2026 to 30% in 2028. If spending on next-generation model training continues to outpace revenue growth, labs will struggle to generate operating leverage. R&D; expenses and capitalization policies must also be monitored to ensure that improving gross margins are not masking broader cost pressure.
The third is business mix. A rising share of enterprise and direct API revenue generally supports paid-inference profitability. If subsidized subscriptions account for too much of the mix, heavy users may erode gross margins. Investors should track API revenue, subscription revenue, usage by paying customers, and changes to plan limits in parallel.
The fourth is the alignment between cloud providers’ AI revenue and capital expenditure. Even with rapid revenue growth, returns on capital may decline if depreciation, lease obligations, and power commitments grow faster. Data-center utilization, transfers from construction in progress to fixed assets, depreciation periods, and operating cash flow can test whether a 35%—45% operating margin is genuinely creating value.
The fifth is infrastructure migration. Commissioning schedules, minimum purchase commitments, financing costs, and actual utilization for dedicated AI projects will determine whether 2028 becomes an inflection point. Delays would prolong hyperscalers’ period of elevated market share. On-time commissioning with inadequate utilization would shift risk to labs and financing providers. Only sustained high utilization would place meaningful pressure on hyperscalers’ share and bargaining power.
The framework’s greatest value is that it translates model competition into a verifiable P&L; chain. Customer spending first enters through subscriptions or APIs and is then allocated between labs and cloud providers. Inference efficiency determines contribution profit, training investment determines operating leverage, and capital expenditure and utilization determine cash returns. As disclosures become more transparent, the market will increasingly shift from narratives about model capability toward comparisons of profit and returns on capital.









