AI Liquid Cooling Deep Dive: 108 Cold Plates, CDUs, and Rubin Fanless Racks: Who Can Defend the US$6-7B Cold Plate Value Pool?
目录
Too Long; Didn't Read
1. Liquid Cooling Is Not a Thermal Accessory; It Is a Constraint on Whether AI Racks Can Be Powered On
2. 108 Cold Plates Explain Why Liquid Cooling Starts With Unit Volume
III. Cold Plates Are a “High-Growth, Low-Service-Attachment” Hardware Business
IV. CDUs Are the “Control Point” of Liquid-Cooling Systems
V. Rubin Fanless Racks Turn “Liquid Cooling as an Option” into the “Platform Default”
6. Supply Chain Ranking: Start with Platform Position, Then Product Form Factor
7. Cold-Plate Model: The Market Is Sizable, But Do Not Misread US$6-7 Billion as Full-Chain Profit
8. Why Air Cooling Will Not Disappear, But Liquid Cooling Will Capture the Core Value
IX. Two-Phase DTC and Direct-to-Die: The Next Liquid-Cooling Reshuffle
X. Domestic Opportunities: Do Not Just Look for “Liquid-Cooling Concepts”; Look at AI Mainline Exposure
XI. Falsification Checklist: When This Theme Should Cool Down
12. What to Track Over the Next Four Quarters
13. Cold-Plate Manufacturing Barriers: The Real Challenge Is Passing “Thermal, Flow, Pressure, and Assembly” at the Same Time
14. Facility-Side Value: Liquid Cooling Pushes Data Centers from “Buying Equipment” to “Buying Availability”
XV. Valuation Framework: Do Not Apply the Same Liquid-Cooling Multiple to Every Layer
XVI. The Real High-Score Answer in the Supply Chain: Moving from Components to Platform Co-Development
XVII. Three Worldviews: What Is Liquid Cooling Really Re-Rating?
XVIII. Two Counterintuitive Points: The Hotter Liquid Cooling Gets, the More Price and Responsibility Matter
XIX. Investment Conclusion: Buy Beta in Cold Plates, Buy Alpha in CDUs and Systems
本内容基于公开资料和研报数据整理,不构成任何投资建议,不代表任何个人观点,仅供学习参考,请理性阅读
As AI racks move from air cooling to liquid cooling, the real change is a redistribution of responsibilities across servers, power delivery, thermal management, and operations. Cold plates will capture volume first; CDUs and rack integration are more likely to defend margins. The next round will hinge on Rubin fanless racks, OCP standards, two-phase cooling, and the pace of customer adoption.
Too Long; Didn't Read
Cold plates get volume first; profits depend on CDUs. A GB200 NVL72 rack contains 72 Blackwell GPUs and 36 CPUs, corresponding to 108 cold plates. As long as AI racks continue to increase power density, cold plates are an unavoidable high-unit-count component in direct-to-chip liquid cooling. The issue is that cold plates look more like precision hardware: service revenue is limited, replacement is mostly handled by server OEMs, and the current US$150-300 ASP plus design premium can support near-term growth. Over the medium term, however, margins are more likely to be compressed by mass production, customer bargaining power, and standardization.
The cold plate pool is about US$6-7B by 2030. Based on incremental GW, liquid-cooling penetration, GPU count, and cold plate value per GPU, the cold plate market can rise from roughly US$2-3B today to about US$6-7B by 2030, implying a four-year CAGR of roughly 20%-30%. The most sensitive variables are not cold plate unit prices, but incremental AI data center GW and liquid-cooling penetration. Once liquid cooling becomes the default option for most new capacity after Rubin, cold plate demand will first show up as unit elasticity.
The technical barrier is concentrated in heat flux. Current cold plates typically support 2-5kW of single-chip heat dissipation, with future designs moving above 5-10kW. Frontier products are already demonstrating 15kW capability. Heat flux targets are moving from 100-130W/cm2 to above 150W/cm2. Pressure drop is usually 15-25psi, while high-performance designs can exceed 30psi. A cold plate is not a simple copper plate. Microchannels, thermal interface materials, pressure loss, and hotspot-directed design together determine whether it can keep up with GPU TDP.
CDUs are more of a systems business than cold plates. CDUs manage flow, pressure, heat exchange, and coupling between two water loops. If a CDU fails, it can affect very high-value AI assets across multiple racks, so customers do not only look at unit price. Single-unit capacity above 2MW, approach temperature around 30 degrees Celsius, OCP compliance, visibility into NVIDIA's roadmap, and field service capability determine whether players such as Vertiv, nVent, Boyd, Motivair, and Delta Electronics can move from "selling equipment" to "selling reliability."
The Taiwan supply chain captures Rubin upside first. Delta Electronics enters through power supplies, CDUs, cold plate modules, and 800VDC racks. Auras Technology, Shuanghong Technology, Fositek, and Sunon are more focused on thermal components and modules. Among A-share companies, Envicool, KSTAR, Kehua Data, Sanhua Intelligent Controls, Lingyi iTech, and BYD Electronics map respectively to opportunities in temperature-control systems, infrastructure, pump/valve and thermal-management components, and structural parts. Ranking should not only ask whether a company has liquid-cooling products, but whether it has entered AI server platforms, has customer validation, and can deliver turnkey solutions.
The real risks are standardization and technology substitution. Single-phase DTC is the main line for 2026-2028, but two-phase DTC, silicon-etched cooling, and direct-to-die will change the form of cold plates and could even weaken the relevance of traditional cold plates over the longer term. Investors should track four numbers: the share of new capacity using liquid cooling, cold plate count and ASP per rack, single-CDU MW capacity and OCP certification, and whether Rubin and subsequent platforms truly move toward fanless racks.
1. Liquid Cooling Is Not a Thermal Accessory; It Is a Constraint on Whether AI Racks Can Be Powered On
The easiest thing to underestimate about AI liquid cooling is that it looks like a "server peripheral," but in substance it has already become part of AI data center delivery speed. As GPU power consumption rises and rack density increases, air cooling does not suddenly disappear; it moves from protagonist to supporting role. Low-density areas, network racks, and some auxiliary equipment still need air cooling, but the main heat from GPUs/CPUs must be removed more directly from the chip surface.
This change creates two investment questions. The first is volume: how many cold plates, CDUs, pipes, sensors, pumps, and valves does each rack need? The second is value: which segments merely follow shipments, and which can retain profits for longer because of system complexity, service responsibility, and customer validation?
The answer is uneven. Cold plates will show volume first, because one GB200 NVL72 rack requires 108 cold plates. CDUs and liquid-cooled rack integration ship in fewer units than cold plates, but they assume responsibility for flow, pressure, heat exchange, leakage, redundancy, and on-site maintenance, making them closer to a "system reliability" business. By Rubin and subsequent platforms, liquid cooling shifts from an "optional upgrade" to a "platform assumption." At that point, thermal management is no longer just about lowering temperature; it determines whether racks can be deployed on schedule.
Rubin Rack Revaluation: GPU Share Falls to 51%, AI Hardware Value Is Migrating Toward Memory, PCBs, Power, and Liquid Cooling; Who Is Capturing the Incremental AI Hardware Opportunity
If liquid cooling is understood only as "water cooling replacing air cooling," the most important pricing logic is missed. An AI rack is not an enlarged version of a consumer-electronics heat sink; it is a safety system for high-value assets. Cooling failure is not simply a slightly higher temperature. It can cause rack-level throttling, downtime, repair, or even damage. Customers are willing to pay for reliability, delivery experience, platform certification, and service networks. This is also why CDUs, rack-level solutions, and suppliers that enter NVIDIA/cloud-provider roadmaps are more likely to defend value than single-component contract manufacturers.
2. 108 Cold Plates Explain Why Liquid Cooling Starts With Unit Volume
AI racks such as GB200 NVL72 turn cold plate demand into a very intuitive multiplication exercise: 18 compute trays, each with 4 Blackwell GPUs and 2 CPUs, for a total of 72 GPUs and 36 CPUs per rack. Each GPU/CPU needs one cold plate, resulting in 108 cold plates. This is more unforgiving than the old framework of "one server, one cooling system," because both the number of heat sources and the power consumption of each heat source in an AI rack are rising.
The role of cold plates in direct-to-chip liquid cooling is straightforward: they are the only components attached to the chip. Coolant exits the CDU, enters the cold plate's internal microchannels through a manifold, the cold plate absorbs heat from the GPU/CPU surface through thermal interface material, and the warmed liquid returns to the CDU. The CDU then transfers the heat through a heat exchanger to the facility-side water loop or to dry coolers/chillers. In this chain, the cold plate is responsible for "removing heat from the chip," while the CDU is responsible for "keeping the liquid circulating in a stable, controllable, and maintainable way."
The near-term certainty of cold plates comes from three variables. First, liquid-cooling penetration is spreading from selected high-density scenarios to more new AI capacity. Depending on the definition, liquid cooling currently accounts for roughly 20%-40% of newly added capacity and is moving toward majority adoption around 2030. Second, GPU TDP continues to rise, requiring stronger heat dissipation per chip. Third, cold plates map almost one-for-one to chip count. Unless the architecture changes fundamentally, more GPUs mean more cold plates.
But "high unit count" does not equal a "wide moat." The value of cold plates lies mainly in design: how microchannels are routed, how thermal interface materials are fitted, how pressure drop is controlled, and how hotspots are addressed directionally. Manufacturing can be expanded, and customers may push for multi-sourcing and standardization. Once TDP growth slows and platform designs stabilize, the innovation premium for cold plates will face pressure faster than that for CDUs.
III. Cold Plates Are a “High-Growth, Low-Service-Attachment” Hardware Business
In the near term, cold plates look like a good business: demand is visible, unit volumes are high, product upgrades are rapid, and customers have to use them. But this is not inherently a durable, high-margin service business. Cold plates are usually designed around the server lifecycle. Failure rates are low, but once leakage, blockage, deposition, or thermal-contact issues occur, the practical remedy is usually not on-site repair of the cold plate itself, but spare-part replacement, reassembly, and testing by the server OEM.
This is completely different from CDUs. A CDU contains pumps, heat exchangers, valves, sensors, control software, and redundancy design. It has a longer life, and failures can affect multiple racks. Service contracts and maintenance capability can become profit sources. Cold-plate vendors mainly earn hardware gross margin, with limited service revenue; if product failure rates exceed contractual thresholds, the cold-plate vendor may instead be forced to bear replacement responsibility. That is not attractive service revenue, but quality cost.
Cold-plate valuation logic therefore needs to be split into two stages. From 2026 to 2028, cold plates remain in the window of “rapidly rising GPU TDP + higher liquid-cooling penetration + platform transitions,” where design and delivery capability can both command premiums. After 2029, if single-phase DTC designs converge, cold plates are more likely to become a competition around scaled manufacturing, yield, cost, and customer share.
This does not mean cold plates have no barriers. True high-end cold plates are not simple CNC-machined parts, but a combination of thermal design, materials, machining precision, and system matching. If thermal interface materials are not handled well, local hot spots can form. The denser the microchannels, the higher the cooling efficiency, but pressure drop and blockage risk also rise. Copper, aluminum, coatings, sealing, welding, and cleaning processes all affect long-term reliability. The barriers are real, but they are more “design barriers during platform transitions” than “system barriers from sustained service lock-in.”
IV. CDUs Are the “Control Point” of Liquid-Cooling Systems
If the cold plate is the heat-absorption end, the CDU is the heart of the liquid-cooling system. It controls coolant flow, pressure, temperature, heat exchange, and redundancy, connecting the server-side technology cooling system with the facility-side water loop. Cold-plate failure usually affects a tray or certain chips; CDU failure can affect a row, a group, or even multiple racks of high-value assets. Customers therefore will not choose CDUs based only on price.
At least five factors matter when assessing CDU suppliers. First, whether single-unit capacity can reach above 2MW, rather than merely using skid stitching to assemble scale. Second, whether approach temperature can remain stable at a high specification level, instead of producing an attractive number only under specific operating conditions. Third, whether the supplier is inside NVIDIA’s power/cooling ecosystem and can see platform power and cooling roadmaps early. Fourth, whether it has OCP-compliant products, especially engineering endorsement from open specifications such as Google Deschutes. Fifth, whether it has a global delivery and service network, because liquid cooling does not end once equipment is sold.
Vertiv, nVent, Boyd, Motivair, and other players have advantages because they are inside key ecosystems and cover rack-level, row-level, and peripheral CDUs. Eaton’s acquisition of Boyd and Schneider Electric’s move to fill its liquid-cooling gap through Motivair show that traditional electrical and building-equipment giants are using M&A; to address liquid-cooling shortcomings. Carrier and Johnson Controls have chiller and HVAC foundations, but their priority and product specifications in CDUs need to be assessed separately. HVAC leadership should not be simply equated with AI liquid-cooling leadership.
AI Materials and Power Supply Ledger Update: How $1.7 Trillion of Capex Reprices Copper, Fiberglass, HVDC, and Liquid Cooling
Another key point about CDUs is that they put liquid cooling and power distribution onto the same diagram. The denser AI racks become, the less possible it is to optimize the power path, busbars, BBUs, PDUs, 800VDC, CDUs, cold plates, and on-site construction separately. Delta Electronics is a typical example: it is not just selling one cold plate or one power supply, but combining AI server power supplies, 800VDC racks, CDUs, and cold-plate modules. For cloud customers, this turnkey capability can reduce integration risk; for suppliers, system capability can turn single-component profit into platform share.
V. Rubin Fanless Racks Turn “Liquid Cooling as an Option” into the “Platform Default”
Rubin’s most important signal is not that a particular cold-plate vendor wins more orders, but that NVIDIA has pushed the cooling assumption for its next-generation platform toward a more complete liquid-cooling architecture. A so-called fanless rack essentially moves heat from inside the server to the liquid loop and facility side, so fans no longer take primary responsibility for GPU/CPU cooling. This will change the profit distribution across the supply chain.
First, cold-plate specifications continue to rise. GPU hot spots are more concentrated, heat flux is higher, and cooling power per GPU is larger. Cold plates cannot merely absorb heat uniformly; they must be customized around hot spots, packaging, stress, and microchannels. Second, CDUs are upgraded from “cooling the liquid” into core components of platform stability, with flow, pressure, redundancy, and sensor data entering the operations loop. Third, the boundary between racks and the facility side becomes blurred: liquid cooling is both an in-rack system and part of the data-center water loop, chillers, dry coolers, and power layout.
Delta Electronics is a representative case. Its liquid-cooling design is moving from liquid-to-air to liquid-to-liquid, and from sidecar CDUs to in-row CDUs. Supply-chain indications suggest Amazon may adopt a self-developed in-row heat exchanger, while Google is advancing its Project Deschutes liquid-cooling solution. Delta Electronics is also viewed as one of the main cold-plate module suppliers for the Vera Rubin platform and is investing in areas such as microchannel lids and two-phase cooling. For investors, this shows that customers do not want a bill of materials; they want system capability that can evolve with the platform.
This also explains why Taiwan’s thermal supply chain has attracted more market attention in the early stage of liquid cooling. Asia Vital Components, Auras Technology, Fositek, Sunon, and other companies have long iterated around server cooling, structural components, hinges/mechanisms, and fan/thermal modules. After entering AI server platforms, platform transitions will amplify revenue elasticity. The differences among companies are: who is closer to NVIDIA/cloud-customer platforms, who has cold-plate module or liquid-cooling system capability, and who is merely following the ASP uplift in traditional thermal solutions.
NVIDIA GTC Deep Dive: From GPU Racks to Token Factories, NVIDIA Brings CPU, Storage, Networking, and Power into the Token Economy, Repricing Full-Stack AI Infrastructure
6. Supply Chain Ranking: Start with Platform Position, Then Product Form Factor
The liquid-cooling supply chain should not be ranked by “who says they have liquid-cooling products.” Four questions matter: whether the company is inside the main AI server platform, whether it can expand from cold plates into CDUs or rack integration, whether it has customer validation and capacity delivery capability, and whether it can withstand pricing pressure after standardization.
In the overseas chain, Vertiv’s advantage lies in integrated data-center power and thermal-management solutions; nVent’s strength is in electrical connectivity, racks, and liquid-cooling components; Boyd has gained stronger resources through its thermal-management history and the Eaton platform; and Motivair fills Schneider Electric’s liquid-cooling capability after entering its system. CoolIT is more focused on cold plates and liquid-cooling components, with clear technical differentiation, but its CDU service attributes and system coverage should be evaluated separately from the categories above. Two-phase cooling companies such as Accelsius and ZutaCore represent longer-term technology shifts and look more like technology options in the near term.
In the Taiwan chain, Delta Electronics looks most like a platform company combining “power + liquid cooling + racks.” Its liquid-cooling revenue growth and data-center revenue mix are both worth tracking. Asia Vital Components and Fositek are more likely to show earnings leverage as AI liquid cooling diffuses. Auras Technology and Sunon also benefit from server thermal upgrades, but the key is specific platform exposure, customer mix, and gross-margin recovery. The core issue here is not “who has risen the most,” but who can move from single thermal components into modules or systems.
The A-share and Hong Kong stock chains require more caution. Envicool, KSTAR, and Kehua Data are closer to data-center temperature control, UPS/power infrastructure, and liquid-cooling systems. Sanhua Intelligent Controls has thermal-management components and pump/valve capabilities, while Lingyi iTech and BYD Electronic are more focused on structural parts, thermal-management modules, and manufacturing capability. Their opportunities are not all at the same layer: temperature-control system companies should be assessed on data-center projects and CDU/cooling-source capabilities; precision manufacturers on whether they enter AI server customer BOMs; and automotive thermal-management suppliers on whether they can migrate pump/valve, heat-exchange, and control experience into data centers.
These companies should not be reduced to a simple “liquid-cooling concept stock” table. Liquid-cooling value is layered along the rack. The closer a company is to platform definition, system integration, and operations responsibility, the more likely it is to generate durable profit. The closer it is to replaceable hardware, the more it must rely on share, yield, and cost to defend returns.
7. Cold-Plate Model: The Market Is Sizable, But Do Not Misread US$6-7 Billion as Full-Chain Profit
The estimate of a US$6-7 billion cold-plate market by 2030 looks attractive, but it is not the entire liquid-cooling market, nor does it mean every participant can capture high profit. The estimate mainly starts from incremental GW, then multiplies by liquid-cooling penetration, GPU count, the number of cold plates per GPU, and ASP. It reasonably reflects the main logic: more AI capacity, more GPUs, more liquid cooling, and therefore more cold plates.
Three sensitivity variables deserve the most attention. First, incremental GW. If AI data-center investment continues to be revised upward, the cold-plate TAM expands passively; if power, land, grid connection, and capex pacing slow, cold-plate revenue will also be deferred. Second, liquid-cooling penetration. Whether Rubin fanless racks can turn into large-scale shipments on schedule will determine how quickly liquid cooling moves from a few high-density scenarios to most new capacity. Third, ASP. Higher GPU TDP will lift cold-plate value per unit, but design stabilization and multi-sourcing strategies will push down unit prices.
More importantly, cold-plate TAM and company profit are not the same thing. Cold plates are high-volume, engineering-intensive hardware with limited attached services. In good years, platform upgrades and tight supply can generate excess profit, but over the long run the category still faces customer multi-sourcing, contract-manufacturing capacity expansion, standardization, and technology substitution. The reasonable expectation for cold-plate makers is not permanently high profit, but to capture share quickly during platform-upgrade windows, expand capacity, improve yield, and extend into modules or system integration.
8. Why Air Cooling Will Not Disappear, But Liquid Cooling Will Capture the Core Value
AI data centers will not become entirely fanless overnight. Low-density general-purpose servers, networking equipment, storage equipment, some auxiliary cooling, and room airflow management will still require air cooling and air-side treatment. The real change is that the main GPU/CPU heat sources can no longer rely on air to remove heat directly. Air cooling moves from “chip-level primary cooling” to “facility-level auxiliary cooling.”
This shift will make liquid cooling and traditional HVAC collaborators rather than simple substitutes. CDUs transfer heat from the server side to the facility water loop, after which chillers, dry coolers, pumps, valves, sensors, and building controls are still needed. HVAC companies, data-center infrastructure companies, and liquid-cooling specialists meet at different boundaries. Whoever can define the interface is more likely to win projects.
For cloud vendors, the difficulty of liquid cooling is not just equipment procurement, but the pace of retrofit. Brownfield data centers must consider existing building water loops, cooling sources, floor loading, leakage risk, and construction windows. Greenfield data centers can be designed from the start around row-level CDUs, liquid-cooled racks, and facility water loops. In the near term, liquid-to-air solutions deploy faster; over the long run, liquid-to-liquid is more efficient. Customers will choose based on site conditions and platform cadence.
This is also why the liquid-cooling value chain will produce two types of winners. One is the “fast-deployment winner,” able to use sidecar CDUs, liquid-cooled racks, and hybrid air/liquid solutions to quickly meet retrofit demand in existing data centers. The other is the “platform winner,” able to follow Rubin, 800VDC, Deschutes, and two-phase DTC toward longer-term architectures. The former is about delivery speed; the latter is about roadmap position.
800V DC Reshapes Data-Center Power: Who Monetizes First, Texas Instruments or onsemi, and How GaN, SiC, and the AI Power Tree Should Be Re-Rated
IX. Two-Phase DTC and Direct-to-Die: The Next Liquid-Cooling Reshuffle
Single-phase direct-to-chip is currently the most mature and easiest-to-scale liquid-cooling path. The coolant remains liquid, absorbs heat through a cold plate, and then returns to the CDU for heat exchange. Its advantages are strong compatibility, relatively low engineering risk, and the ability to connect with existing data-center facilities. The issue is that as rack power continues to rise, single-phase DTC requires higher flow rates, more complex cold plates, lower thermal resistance, and stronger pump power.
The logic of two-phase DTC is to let the cooling medium undergo a phase change during heat absorption, using latent heat of vaporization to remove more heat. It may improve cooling capacity, but it also brings issues around medium selection, pressure, sealing, control, maintenance, and customer acceptance. At this stage, it looks more like a technology reserve for the next phase than the main path that will fully replace single-phase DTC in 2026. For suppliers, early positioning in two-phase DTC provides a technology option, but near-term revenue still mainly comes from single-phase DTC and hybrid air-liquid solutions.
Direct-to-die and silicon-etched cooling are more aggressive. They integrate cooling deeper into the chip package or silicon structure, which in theory could reduce the value of traditional cold plates or even change the form factor of cold plates. True commercialization requires joint acceptance by chip vendors, packaging houses, server OEMs, and data-center operators, so the cycle will not be short. But this is exactly the long-term risk for cold plates: cold plates are essential today, but may be redefined by package-level cooling in the future.
The CDU’s position in this transition is more stable. Two-phase DTC requires more complex control, and direct-to-die still needs to move heat from the server to the facility side. Liquid circulation, heat exchange, monitoring, redundancy, and maintenance will not disappear. Traditional cold plates may be replaced, but liquid-cooling control systems are harder to eliminate.
X. Domestic Opportunities: Do Not Just Look for “Liquid-Cooling Concepts”; Look at AI Mainline Exposure
There are many domestic liquid-cooling companies, but opportunities truly tied to AI racks require another layer of screening. Traditional data-center liquid cooling, energy-storage thermal management, telecom thermal management, industrial cooling, and AI server liquid cooling are not the same thing. AI racks require higher heat flux density, stricter leakage control, shorter delivery cycles, stronger customer validation, and higher asset risk. Not all liquid-cooling revenue should be treated as AI liquid-cooling revenue.
Envicool’s advantages lie in data-center thermal management and liquid-cooling system experience. The key is overseas AI customers, CDU/liquid-cooled rack capability, and execution on high-power projects. KSTAR and Kehua Data are closer to UPS, power supply, and data-center infrastructure; their liquid-cooling opportunities should be assessed together with power distribution systems. Sanhua Intelligent Controls is migrating from automotive thermal management to data centers, and needs to prove whether its pumps, valves, heat exchange, and system control can enter AI server and facility-side BOMs. For precision manufacturing companies such as Lingyi iTech and BYD Electronic, the key is whether cold plates, structural parts, and thermal modules enter mainstream AI server platforms, not a generic claim of “thermal management capability.”
A more direct judgment method is to place the company into the rack diagram: does it make the cold plate on the chip, the manifold inside the server, in-rack piping, row-level CDUs, facility water loops, chillers, pumps and valves, control software, or project integration? Different positions imply completely different valuation logic. Cold plates and structural parts depend on platform share and yield; CDUs depend on system reliability and service; cold sources and facility-side equipment depend on the data-center construction cycle; pumps and valves depend on certification and value per rack.
There is another practical issue in the A-share chain: supply-chain certification for overseas AI server main platforms is slower, and customers are usually unwilling to take risks by switching key thermal-management suppliers on high-value racks. For domestic companies to move from “theme” to “earnings,” investors need to see clear customers, platforms, product forms, and revenue mix, rather than only prototypes, order rumors, or conceptual descriptions.
XI. Falsification Checklist: When This Theme Should Cool Down
Liquid cooling is not risk-free. Precisely because market expectations are hot, falsification indicators need to be written down in advance. The first falsification point is a slowdown in new AI data-center capacity. If power, grid connection, land, capex, or cloud-vendor demand changes lead to downward revisions in new GW additions, the shipment elasticity of cold plates and CDUs will be delayed.
The second falsification point is a slower-than-expected Rubin liquid-cooling pace. The market treats fanless racks as the long-term default for liquid cooling. If the Rubin platform is delayed, adopts a more conservative hybrid air-liquid solution, or customer migration falls short of expectations, cold-plate ASPs and CDU orders may both decline.
The third falsification point is an overly rapid decline in cold-plate ASPs. The current design premium for cold plates comes from rapid platform iteration. Once multi-sourcing matures, customers unify specifications, and ODM yields improve, pricing may compress earlier than the market expects. Volume can still grow, but profit will be hit first.
The fourth falsification point is leakage, blockage, or maintenance incidents. In high-value AI racks, reliability events are the biggest concern for liquid cooling. One serious incident could change the pace of customer certification. Suppliers with only lab parameters and no scaled field record should receive a valuation discount.
The fifth falsification point is a technology-path leap. If two-phase DTC, direct-to-die, or package-level cooling commercializes ahead of schedule, it would compress the long-term space for traditional cold plates. This risk should not be exaggerated in the short term, but it must be reflected in valuation and holding period.
12. What to Track Over the Next Four Quarters
The most useful tracking framework for liquid cooling is not whether there are orders, but a breakdown by platform, product, pricing, and reliability. For platforms, track whether Rubin and GB300/Rubin Ultra-related racks ship on schedule. For products, track the number of cold plates per rack, cold-plate power ratings, CDU capacity, and row-level deployment methods. For pricing, track cold-plate ASPs, liquid-cooling module gross margins, and CDU service revenue. For reliability, track leaks, clogging, repairs, and on-site maintenance.
For the Taiwanese thermal-management chain, the focus should be revenue elasticity and gross margin during the platform transition from 2Q26 to 2027. If cold-plate/liquid-cooling module revenue grows quickly but gross margin does not improve, the company is more of a manufacturing contractor. Only if revenue growth comes with product-mix and gross-margin improvement does it indicate a design premium. For platform-type companies such as Delta Electronics, also track the share of liquid-cooling revenue, the share of data-center revenue, and whether 800VDC and CDU solutions scale together.
For overseas system companies, focus on OCP certification, the NVIDIA ecosystem, shipments of CDUs above 2MW, service contracts, and M&A; integration. Vertiv, nVent, Eaton/Boyd, and Schneider/Motivair should not be judged only by single-quarter orders; the key is whether they are incorporating liquid cooling into long-term data-center infrastructure solutions.
For the A-share chain, focus on the true revenue scope. Liquid-cooling prototypes, energy-storage thermal management, traditional IDC projects, and AI server liquid cooling cannot be mixed together. What can truly lift valuation is exposure to overseas AI customers, mainstream server platforms, clear BOM positions in CDUs/cold plates/pumps/valves, and sustained shipments.
In-Depth 13F Analysis of 327 Major Institutions | NVIDIA and AVGO Remain Core AI Holdings, While Power, Thermal Management, Ethernet, Optical Components, and Memory Become Marginal Allocation Increases
13. Cold-Plate Manufacturing Barriers: The Real Challenge Is Passing “Thermal, Flow, Pressure, and Assembly” at the Same Time
The long-term debate around cold plates is whether they will become commoditized. This cannot be answered simply with “yes” or “no.” Low-end cold plates will commoditize. High-end AI cold plates will not commoditize too quickly during platform transitions. But once platforms stabilize, customer specifications converge, and multi-sourcing matures, manufacturing cost and yield will become more important variables.
The first challenge in cold-plate design is thermal. GPU heat sources are not uniform. Tensor cores, HBM, package edges, and power-delivery areas have different thermal distributions. Internal cold-plate channels cannot simply pursue average temperature; they must address local hotspots. The more aggressively hotspots are handled, the more complex the microchannels, and the lower the thermal resistance, but processing difficulty and clogging risk also increase.
The second challenge is flow. Faster coolant flow is not always better; there are trade-offs among velocity, flow rate, pressure drop, and pump power consumption. The denser the microchannels, the larger the contact area and the better the heat exchange, but pressure drop will rise, requiring the CDU to use higher pump power and stronger controls to maintain system stability. If a cold-plate maker only optimizes its own thermal parameters and does not conduct system matching with the CDU, manifold, and server OEM, real rack-level performance will be discounted.












