< img src="https://mc.yandex.ru/watch/103289485" style="position:absolute; left:-9999px;" alt="" />

How to Choose the Right CDU for AI and HPC Racks: A Selection Guide

A coolant distribution unit in an AI data center with glowing liquid pipes and digital telemetry overlays

Key takeaway: Choosing a coolant distribution unit (CDU) is less about picking the biggest pump and more about matching the unit to your density, redundancy, and facility constraints. Work through the needs assessment before looking at specs, then score candidates against the criteria that map to your workload's availability target.

Air cooling stops scaling cleanly once racks pass roughly 20–40 kW, which is where most AI and HPC workloads now sit. At that point the coolant distribution unit (CDU) becomes the piece that turns a loose bundle of pipes into an engineered loop: it pumps coolant, exchanges heat, filters and conditions the fluid, and isolates the IT loop from the facility water system (FWS). Pick the wrong unit and you get marginal thermal headroom, awkward maintenance, or a redundancy model that does not match your availability goal.

This guide walks through a repeatable selection process. It starts with the questions you answer before you open a spec sheet, then breaks down the architecture choice, the thermal and hydraulic criteria, reliability planning, facility integration, and the procurement and acceptance steps that turn a shortlist into a delivered, commissioned unit.

The single most common mistake in CDU selection is leading with published kW capacity and working backward. Capacity only matters once you have pinned down a few structural facts about the deployment.

Rack and cluster density forecast. You need the design heat load per rack and per cluster over the next 12 to 24 months, including the growth you expect from future GPU generations. Size the CDU to the most power-dense racks and the total load per compute zone, because undersizing the headroom is what causes thermal throttling later.

Loop type. A liquid-to-liquid (L2L) CDU exchanges heat against a facility water or chilled-water loop and is generally the higher-capacity, more efficient option where water is available. A liquid-to-air (L2A) CDU rejects heat to ambient air and is easier to deploy where water is scarce or not plumbed, at the cost of lower capacity and higher energy use. Choose by water availability first, not by a vendor preference.

Secondary loop temperature target. Decide your ΔT across the technology cooling system and any warm-water target before you compare units. Locking this first prevents you from comparing units at different operating points.

Integration map. A CDU touches the facility water system, the electrical feeds, and the BMS or DCIM. Map these early so the unit you choose can actually talk to the systems you already run.

Two adjacent questions also shape everything downstream. First, is the deployment new-build, retrofit, or a mix? Retrofit sites face structural, space, and piping constraints that can rule out some form factors. Second, who owns the water chemistry: your facilities team or the vendor? That ownership decision comes up again under water quality.

Match the architecture to the deployment

CDUs come in three common physical forms, and the architecture choice drives almost every other criterion you will score.

Architecture

Typical fit

Redundancy model

Best when

Rack-mounted CDU

Ultra-dense single racks, tightest loop control

Redundant pumps/feeds within each unit; isolation limits failure to one rack

Rack-local design, high density, tight thermal control per cabinet

Row-level CDU

Repeating AI pods, centralized maintenance

One unit serves multiple racks; N+1 at row/pod level

Larger AI pods, efficient scaling across many racks

Inline / distributed CDU

Flexible modular layouts, balanced redundancy

Internal redundancy; system resilience depends on how many parallel units are deployed

Flexible modular growth near the load

Rack-mounted units give the shortest coolant path, which usually means the lowest approach temperature for a single dense rack, but scaling them rack by rack raises cost and service burden. Row-level units centralize capacity and scale cleanly across pods, which is why they dominate larger AI deployments, though longer row piping adds a little hydraulic loss. Inline units sit in between, trading per-unit capacity for placement flexibility near the load.

Your answer here is not a universal winner. If the deployment is a handful of very dense, localized racks, rack-mounted cooling wins on thermal control. If it is a repeating AI pod, row-level usually wins on economics and serviceability. If flexibility is the priority, inline is the middle ground. This decision also depends on how high you intend to push density: at the 20 to 40 kW per rack inflection band, direct-to-chip liquid cooling becomes the primary path, which in turn drives the CDU form factor you need.

Score the thermal and hydraulic criteria

With the form factor settled, you move to the criteria that determine whether a unit can actually move the heat at the conditions your equipment needs.

Flow rate. A practical starting point in current AI deployments is around 1.5 to 2.0 liters per minute per kilowatt of heat. Flow must match your ΔT target and the cold-plate requirements, not just the nameplate capacity. Coolnetpower’s CDU sizing guide walks through the full mass-flow and pressure-drop method with a worked example, using the Q = m × cp × ΔT relationship.

Pump head. The pump has to overcome the pressure drop across cold plates, manifolds, and piping with margin. Under-specify this and flow collapses at the far end of the row; over-specify and you waste pump energy. Ask for the pump curve at realistic operating points, not just a headline pressure figure.

Approach temperature. This is the temperature difference between the cooling medium and the process fluid leaving the heat exchanger. Lower approach temperature gives more thermal headroom and improves economizer effectiveness, but it is not a universal number to chase. Size it to the actual heat exchanger, flow, and fouling assumptions rather than copying a competitor’s target.

Turndown and stability. AI loads surge and collapse quickly. The unit needs to hold stable supply temperature and flow across rapid load change, which means capacious pump turndown and responsive control, not just peak capacity.

For every one of these, ask for the number at the operating point you actually need, plus a partial-load figure. A unit that looks strong at design load can be weak at the 40% loading where it will spend most of its life.

Plan reliability and redundancy at the right layer

Redundancy is where well-intentioned CDU choices go wrong, usually because reliability is treated as a single switch rather than a layered property. The Uptime Institute’s guidance on close-coupled cooling and reliability is direct on this: the design fails or passes depending on whether removing any single cooling unit pushes a cold aisle past its redundant count. Plan redundancy independently at four levels, not just inside one box.

Within the CDU. Look for redundant pumps, redundant controllers, redundant power feeds, and redundant sensors. N+1 pump redundancy, where one extra pump can carry the design flow after a single pump fails or is pulled for service, is the common baseline. More advanced units use parallel pump sets with automatic failover and near-zero parasitic load from the idle pump.

Across the CDU units. If one CDU serves more racks than the remaining units can carry, you have a single point of failure. Size so that with one unit offline, the rest still deliver design flow and approach temperature at full rack load. Vertiv’s overview of N+1 redundancy in data center cooling makes the same point at the cooling-unit level: N+1 means one spare unit beyond the count required to meet full load.

On the facility side. Redundancy must extend to the distribution loop and the heat-rejection plant. A redundant pump inside a CDU does nothing if the chilled-water source feeding it is a single point.

Lock setpoints to OEM curves, not invented numbers. Approach temperature, supply temperature, and dew-point margin should be tied to the OEM coolant curves and your site’s condensation limits, then validated, rather than set to a generic industry value.

A 2N model, where every component has a mirrored backup, adds resilience but at roughly double the capital cost. Most operators pick N+1 at the pump and unit level and reserve 2N for the most critical or highest-value deployments. Coolnetpower’s detail on CDU sizing and redundancy for AI loads sets out how to choose between N+1 and 2N based on the number of racks served and the cost of downtime.

Cover water quality, filtration, and leak detection

Liquid loops reward discipline, because a small chemistry or contamination issue becomes a system-wide reliability problem over time.

Filtration. The technology cooling system loop generally needs stricter filtration than the facility side. A commonly cited reference point is the filtration guidance from the Open Compute Project’s liquid-cooling work, which specifies around 25 to 50 micrometers on the secondary or TCS side and coarser primary filtration on the facility loop, though you should verify limits against your OEM and coolant spec. Look for differential-pressure monitoring and bypass or serviceable filters so fouling never becomes a single point of failure.

Water chemistry. Define wetted materials, inhibitors, coolant type, and acceptable contamination limits explicitly, because corrosion, fouling, and seal degradation all trace back to compatibility. Strict particulate and ion limits apply to TCS loops; treat them as application-specific and confirm them against OEM requirements rather than assuming generic chilled-water values work.

Leak detection. Layer it. Combine point or cable sensors at connectors and serviceable components with pressure and flow trend monitoring to catch slow leaks or air ingress. Then define response logic so operators know which alarms mean what, what gets isolated, and who responds. Add isolation valves, drip management, and maintainable branches so a leak can be contained without draining the whole loop.

Fit the CDU to your facility and control systems

A CDU that performs on paper but cannot integrate cleanly into your facility will disappoint in operations. Two integration areas carry most of the risk.

Facility water loop. The CDU is the boundary that decouples the facility water system from the technology cooling system. Getting this handshake right matters enough that Coolnetpower has a dedicated guide to CDU facility loop integration that covers piping, isolation, and balancing boundaries.

BMS and DCIM. Treat the CDU as a telemetry and alarming source connected to your BMS or DCIM, but do not let a single supervisory system become the control plane for basic cooling flow. The unit should sustain flow locally even if the supervisory system fails. Monitor the signals that indicate reliability: pump status and speed, flow, differential pressure, supply and return temperatures, filter ΔP, leak zones, and valve position. Clear alarm ownership keeps the facilities team, the white-space operations team, and the vendor from all assuming the next person handled it.

Turn the shortlist into accepted hardware

Selection does not end when you sign the PO. The procurement and acceptance phase is where a good choice becomes a working installation, and it deserves the same rigor as the technical scoring.

Write the criteria into the RFP first. Build a checklist that states the must-have criteria and the red flags separately, so vendors respond against the same standard. Must-haves typically include the flow and approach-temperature numbers at your operating point, the redundancy model, and the integration points. Red flags include capacity claims that collapse under partial load, single points of failure where you asked for N+1, and vague answers on commissioning.

Validate before you accept. Factory acceptance testing (FAT) and, after delivery, integrated systems testing exist to catch problems before they cost you uptime. Test hydraulic performance across the operating range, including low-load and failover cases. Simulate a pump fault, a lost sensor, an isolated branch, or a leak event under representative load, then measure ride-through and temperature stability. Confirm maintainability, too: isolation valves, quick disconnects, and access to filters and pumps should let you work under load where the design claims concurrent maintainability.

Verify the numbers in the field. Do not trust nameplate performance. Measure flow, pressure drop, approach temperature, and pump energy at realistic points during commissioning, and record them against the FAT results.

Red flags that should stop a CDU shortlist

A few warning signs warrant an immediate question or a pass, depending on how critical the deployment is:

  • Capacity quoted only at a favorable operating point with no partial-load figures.

  • Single pump or single controller where your workload requires N+1 or 2N.

  • No clarity on TCS-side filtration or water-chemistry ownership.

  • Vague answers on leak detection, isolation, or alarm ownership.

  • A form factor that cannot fit your space, weight, or piping constraints.

  • No defined commissioning or FAT procedure.

Key takeaway: The deal-breakers are almost never about peak capacity. Redundancy at the right layer, defensible approach temperature at your operating point, real filtration and leak protection, and a defined acceptance test are what separate a CDU that protects AI uptime from one that merely meets a brochure number.

Next steps

Choosing a CDU comes down to answering the needs assessment, then scoring candidates against the criteria that match your availability target and facility constraints. A cheap unit that throttles under an inference burst costs far more than the savings on the order.

If you want the selection criteria turned into a working document, Coolnetpower’s engineering team can supply a CDU selection and RFP checklist tailored to your density and redundancy targets. Start with the high-density coolant distribution units product page to see how the rack, row, and large form factors map to real deployments, then bring your density forecast and loop design to a technical fit call.

FAQ

What does a coolant distribution unit (CDU) do? A CDU pumps coolant, exchanges heat between the facility water system and the technology cooling system, and filters and conditions the fluid so the IT loop stays within safe temperature, flow, and pressure limits. It is the isolation and control boundary that keeps clean, conditioned coolant flowing to cold plates or other liquid-cooled equipment.

How do I choose between a rack-mounted and a row-level CDU? Use rack-mounted units when the design is rack-local, ultra-dense, or needs the tightest per-cabinet thermal control. Use row-level units for repeating AI pods where centralized capacity, easier maintenance, and lower cost across many racks matter more. The choice reflects your density pattern and service model, not a universal best option.

What is a good approach temperature for a CDU? There is no single good value. Lower approach temperature improves headroom and economizer effectiveness, but the right target depends on your heat exchanger, flow, and fouling assumptions. Size approach temperature to your actual operating point and OEM curves instead of chasing a competitor’s number.

Is N+1 or 2N better for a CDU? N+1 adds one spare pump or unit so a single failure or maintenance event does not drop load, which covers most missions. 2N mirrors every component and adds resilience but roughly doubles capital cost. Choose based on how many racks a unit serves and how expensive downtime is for your workload.

Why does water quality matter for a liquid cooling loop? Contaminants, incompatible materials, and poor chemistry drive corrosion, fouling, and seal degradation over time. Stricter TCS-side filtration and defined water chemistry protect the CDU and the cold plates so the loop stays reliable instead of degrading into a maintenance problem.

Facebook
Pinterest
Twitter
LinkedIn

Leave a Reply

Your email address will not be published. Required fields are marked*

About the author

Rajon

Rajon

As a dedicated technical marketing professional in the data center infrastructure and thermal management sector, Rajon specializes in precision cooling and modular systems. Combining engineering logic with data-driven B2B strategies. Through this hands-on industry experience, Rajon translates complex concepts into clear, actionable insights for professionals worldwide.
Tel
Wechat