Data Center

Direct-to-Chip Liquid Cooling: CDU, Cold Plates & Facility Design

Air cooling hit its ceiling. At 40 kW per rack, the volume of cold air required to prevent thermal throttling exceeds what even the most aggressive hot-aisle containment and in-row cooling can deliver. The industry response is not a better fan — it is a fundamentally different heat transfer medium. Water carries heat 3,500 times more efficiently than air by volume. Direct-to-chip liquid cooling exploits that physics advantage by placing copper or aluminum cold plates directly on the hottest components in a server, removing the majority of heat before it ever enters the room.

This guide covers the complete direct liquid cooling (DLC) stack: from the cold plate on the GPU die to the facility chilled water plant that ultimately rejects the heat. If you are planning a high-density deployment, retrofitting an existing facility, or evaluating GPU colocation providers, understanding this infrastructure is essential.

How Direct-to-Chip Cooling Works

The system has four layers, each transferring heat from a hotter medium to a cooler one:

  1. Cold plates: Copper or aluminum blocks with internal microchannels, mounted directly on GPUs, CPUs, and high-power VRMs (voltage regulator modules). Coolant flows through these channels, absorbing heat through conduction. A single NVIDIA H100 SXM5 GPU produces up to 700W of heat — the cold plate captures approximately 95% of that at the die.
  2. Manifolds and hoses: Quick-connect manifolds at the rear of each server distribute coolant from a rack-level supply loop to individual cold plates and collect the heated return flow. Drip-free quick-disconnect fittings allow servers to be removed for maintenance without draining the loop.
  3. Coolant Distribution Unit (CDU): A heat exchanger, pump, reservoir, and control system that sits either in the rack row (in-row CDU) or in a centralized mechanical room. The CDU transfers heat from the server-side coolant loop (typically propylene glycol/water mix) to the facility-side chilled water loop. It also maintains coolant pressure, flow rate, temperature, and quality.
  4. Facility rejection: The facility chilled water system (chillers, cooling towers, or dry coolers) ultimately rejects the heat to the atmosphere. In hot climates, this may require mechanical chillers year-round. In temperate regions, free cooling can handle most of the load, dramatically reducing energy consumption.

The result: 70-80% of total server heat is removed by liquid before it enters the room air. The remaining 20-30% (from DIMMs, NVMe drives, network cards, and fans) is handled by supplemental air cooling, which can be far simpler and cheaper because the heat load is drastically reduced.

Direct Liquid Cooling vs. Immersion Cooling

Both use liquid, but they are architecturally different solutions for different constraints:

Factor Direct-to-Chip (Cold Plate) Immersion Cooling
Form factorStandard 19-inch racksCustom tanks or tubs
Heat removal70-80% via liquid, 20-30% via air100% via dielectric fluid
CoolantWater/glycol mix (conductive — must not leak)Dielectric fluid (non-conductive)
Server compatibilityRequires DLC-ready servers with cold plate mountsAny hardware (submerged as-is)
Maintenance accessStandard rack serviceabilityMust lift hardware from fluid
Leak riskHigher — water near electronicsLower — fluid is non-conductive
Max densityUp to ~150 kW/rackEffectively unlimited
MaturityProduction-proven at hyperscaleGrowing but less standardized
Best forGPU training clusters, enterprise HPCExtreme density, edge, overclocked mining

For most AI and HPC deployments in 2026, direct-to-chip cooling is the practical choice because it preserves standard rack infrastructure, uses commercially available DLC-ready servers from every major OEM, and has a mature supply chain. Immersion excels where maximum density or unique form factors justify the operational complexity.

CDU Architecture: In-Row vs. Centralized

In-Row CDUs

Placed between server racks (consuming one rack slot per 4-8 racks), in-row CDUs keep coolant loops short, minimizing pump energy and leak exposure. Each CDU handles 50-200 kW of heat rejection. This is the preferred architecture for colocation environments where different customers may have different cooling requirements, because each CDU can be independently controlled.

Centralized CDUs

Located in a dedicated mechanical room, centralized CDUs serve entire rows or halls through longer piping runs. They offer economies of scale (fewer, larger units are more efficient per kW) but require more extensive facility plumbing and create a single point of failure if not configured with N+1 redundancy. Best suited for operator-controlled environments where the entire hall runs homogeneous workloads.

CDU Sizing

A CDU must handle the peak liquid-side heat load of the racks it serves. For a row of 10 racks at 80 kW each, the liquid-side load is approximately 10 × 80 × 0.75 = 600 kW (assuming 75% of heat removed by liquid). With N+1 redundancy, you need CDU capacity for 600 kW plus one spare unit. Undersizing the CDU is the most common deployment error — it causes coolant temperature to rise, which throttles GPUs and reduces training throughput.

Facility Plumbing Requirements

Retrofitting or building for DLC requires facility-level infrastructure that does not exist in air-cooled data centers:

  • Supply and return piping: Chilled water supply (typically 7-12°C) and warm return (typically 35-45°C) must be routed from the chiller plant to each CDU location. Pipe diameter depends on flow rate, which depends on total heat load. A 1 MW liquid cooling deployment typically requires 6-inch supply and return mains.
  • Leak detection: Water near live electronics creates a leak-damage risk that air cooling does not have. Leak detection cables must be installed under every rack row, around every CDU, and at every manifold connection. Detection systems should trigger automatic pump shutdown and operator alerts within seconds.
  • Water treatment: Coolant quality (pH, conductivity, biocide levels, particulate count) must be maintained to prevent corrosion, biological growth, and cold plate fouling. Inline filters and periodic testing are standard. Some facilities use closed-loop deionized water systems to minimize contamination risk.
  • Redundancy: Pump failure means cooling failure means thermal shutdown. CDUs should have redundant pumps (N+1 minimum), and critical deployments should have redundant CDUs. Valve isolation enables maintenance on individual CDUs without shutting down the entire cooling loop.
  • Heat rejection upgrade: The facility chiller plant or cooling tower must have capacity for the additional liquid-cooled load. For a facility converting from 8 kW/rack air-cooled to 80 kW/rack liquid-cooled (even partially), the cooling plant capacity requirement can increase by 3-5x. This is often the largest capital expense in a retrofit.

TCO: Liquid Cooling vs. Air Cooling

The economics shift decisively in favor of liquid cooling above 30 kW per rack:

  • Capital cost premium: DLC adds approximately $200-$500 per kW of IT load over air cooling for cold plates, manifolds, CDUs, piping, and leak detection. For a 1 MW deployment, expect $200K-$500K in additional cooling infrastructure.
  • PUE improvement: Air-cooled facilities typically achieve PUE 1.4-1.6. DLC facilities achieve PUE 1.1-1.2. At $0.10/kWh and 1 MW IT load, the difference between PUE 1.5 and 1.15 is approximately $307,000 per year in energy savings.
  • Density advantage: DLC enables 3-5x higher rack density. A workload that needs 100 air-cooled racks at 10 kW each fits in 20-25 liquid-cooled racks at 40-50 kW each. This reduces floor space, cabling, network switches, and management overhead.
  • GPU performance: Liquid-cooled GPUs run at lower junction temperatures, enabling sustained boost clocks and avoiding thermal throttling. For AI training workloads, this can translate to 5-10% higher throughput compared to air-cooled equivalents at the same power budget.

For facilities in the UAE and Middle East where ambient temperatures regularly exceed 40°C, the PUE advantage of liquid cooling is even more pronounced because air-side economizers are rarely viable. DLC with warm-water cooling (inlet temperatures up to 45°C) can operate with dry coolers instead of mechanical chillers for much of the year, dramatically reducing energy costs.

Deployment Checklist

  1. Confirm server compatibility. Verify your chosen server platform supports DLC (OEM cold plate mounts, manifold connections). NVIDIA DGX H100 and B200 ship DLC-ready. Dell, HPE, Supermicro, and Lenovo all offer DLC-ready GPU server SKUs.
  2. Size the CDU. Calculate total liquid-side heat load (total rack power × 0.70-0.80). Add N+1 redundancy. Select CDU models that match.
  3. Design facility plumbing. Route supply/return piping from chiller plant to CDU locations. Specify pipe diameter for required flow rate. Include isolation valves per CDU for maintenance.
  4. Install leak detection. Cable sensors under every rack row and CDU. Connect to BMS (building management system) with automatic pump shutdown on detection.
  5. Verify structural capacity. Liquid-cooled racks with full coolant weigh more than air-cooled equivalents. Confirm floor loading.
  6. Commission and test. Pressure-test all piping before filling with coolant. Run CDUs at full flow before energizing servers. Verify coolant temperature, flow rate, and leak detection at every connection point.

Related Articles

Liquid-Cooled Infrastructure Ready

Rax facilities support direct-to-chip liquid cooling with CDU infrastructure, facility plumbing, and leak detection built in. Deploy your GPU clusters at the density your workloads demand.

Request a Quote