8 min

How to choose cooling for an accelerator rack

Compare accelerator rack cooling for Kazakhstan's climate, including air limits, liquid loops, leak risk, lifecycle cost, and maintenance.

How to choose cooling for an accelerator rack

Air and liquid cooling should not be chosen by fashion or by a city's average summer temperature. For an accelerator rack, the choice depends on the equipment's maximum electrical power, allowable inlet temperature, the room's actual ability to remove heat after one component fails, and whether the operations team can maintain the chosen system.

Kazakhstan's dry air and cold winters provide many hours of economizer operation, but they do not turn an 80 or 120 kW rack into an ordinary server-room load. If the existing room can supply the required volume of clean air without recirculation, an air-cooled design remains sensible. If that airflow cannot pass through the rack, raised floor, aisle, or cooling units, direct liquid cooling becomes an engineering requirement, not an expensive option.

Calculate heat before counting accelerators

The cooling system must remove almost all electrical power used by the IT equipment as heat, so calculate kilowatts per rack at the design load rather than the number of servers or the rated power of one GPU. Include accelerators, processors, memory, network switches, storage, power supplies, and losses inside the rack. Add cooling equipment powered by the same electrical system separately, but do not count its consumption as useful IT load.

The first working document for the project can contain five lines:

P_IT_MAX_KW=
AIR_HEAT_FRACTION=
LIQUID_HEAT_FRACTION=
DESIGN_INLET_C=
REDUNDANCY_MODE=N+1

Take P_IT_MAX_KW from the agreed configuration and power profile, not from the rating of the input breaker. The heat fractions must add up to one. For a fully air-cooled server, AIR_HEAT_FRACTION equals 1. In a hybrid system, cold plates may remove most of the load from the CPUs and GPUs, while fans still remove heat from memory, power converters, storage, and network components. The manufacturer must provide these fractions for the exact configuration.

For the air portion, a heat-balance check is useful: volumetric airflow is approximately Q / (ρ × cp × ΔT). At 100 kW, with air density of 1.2 kg/m³, specific heat of 1.005 kJ/(kg·K), and a 12°C air temperature rise, the result is about 6.9 m³/s, or roughly 24,900 m³/h. This is not the design airflow, only a quick order-of-magnitude check. Filters, altitude, leakage through blanking panels, door resistance, and redundancy will increase the required capacity.

Altitude matters especially in Almaty and at sites in the foothills. At lower air density, the same mass flow requires a greater air volume, and the fans operate closer to their limit. Do not reuse a sea-level calculation without an altitude correction and a check against the fan curves.

NVIDIA's documentation for DGX H100/H200 provides a useful sense of scale: it specifies maximum system power of 10.2 kW, an operating temperature of 5-30°C, relative humidity of 20-80% non-condensing, and airflow of 1,105 CFM at 80% fan PWM. Those figures apply to one specific system. The lesson is not to copy them into another project, but to request the same complete data set from your own supplier.

Kazakhstan's climate helps outside the rack

Kazakhstan's continental climate increases the potential for free cooling, but it also forces designers to account for heat, dust, very dry air, and winter freezing. A July average is almost useless for this task. You need hourly climate data for the specific city, the design maximum dry-bulb temperature, the joint distribution of temperature and humidity, dusty periods, prevailing wind, and the minimum winter temperature.

Kazhydromet publishes the State Climate Cadastre with long-term temperature series and station extremes. In its 2023 annual bulletin, the service recorded a maximum of +46.0°C at the Karatobe station in West Kazakhstan Region. That is not the design value for every site in the country, but it is a strong reason to reject a single "Kazakhstan" climate profile. Astana, Almaty, Aktau, Atyrau, and Shymkent need different design points and control modes.

Dry outdoor air is useful for air-side or water-side economizing when its temperature and moisture content allow compressors to be switched off or unloaded. It also makes evaporative cooling more effective. Free cooling still has operating costs: outdoor air needs filtration, dampers and sensors need maintenance, and an evaporative system consumes water and leaves salts in the circuit.

A direct air-side economizer in a steppe climate cannot be evaluated by temperature alone. Dust raises the pressure drop across filters and reduces airflow, while smoke or urban pollution may disable outdoor-air mode on the worst possible day. The US Department of Energy's guide to energy-efficient data centers recommends accounting for contaminants and humidity and locking out the economizer by dew point; it cites MERV 13 filtration as a common measure against outdoor particles. For a project in Kazakhstan, choose the filter class from a local environmental analysis and the equipment requirements, not from one foreign recommendation.

Very dry air is not a reason to humidify the whole room without control. ASHRAE advises monitoring moisture by dew point because relative humidity changes with temperature across the room. Set humidification according to the permitted range for the specific servers, the air-delivery method, and an electrostatic-risk assessment. An unnecessarily narrow range can make adjacent units humidify and dehumidify in turn, wasting energy and water.

Winter brings different hazards: freezing in outdoor heat exchangers, dry coolers, and idle pipe sections; fine snow at the air intake; cold starts; and condensation during poorly controlled mode changes. For an outdoor liquid circuit, verify glycol concentration, heat tracing, drainage, backup pump power, and restart logic after downtime. Low winter temperatures reduce heat-rejection costs only when the controls can use them safely.

Air cooling is limited by the airflow path

An air-cooled design works for an accelerator rack while the complete path can deliver the design flow to the front and remove hot air without mixing. The cooling units' nominal capacity does not prove that. I have seen rooms with spare refrigeration capacity where top servers overheated because exhaust air took a short path back from the hot aisle to the cold aisle.

Check the entire sequence: cooling-unit outlet, plenum or overhead distribution, perforated tile or grille, cold aisle, rack doors, server fans, hot aisle, and return to the cooling unit. The narrowest point sets the limit for the whole system. If a rack needs 25,000 m³/h, installing another chiller achieves nothing when the grilles and raised floor pass only part of that flow.

Cold-aisle or hot-aisle containment usually helps more than lowering the cooling-unit setpoint. NVIDIA's DGX SuperPOD design guide explains directly that aisle containment prevents hot exhaust from recirculating to server inlets, which artificially raises their temperature and reduces the air's capacity to absorb heat. I agree with that priority: separate the flows and fill empty rack units with blanking panels before changing the setpoint.

Air cooling has clear strengths:

  • there is no water or quick-disconnect coupling inside the server rack;
  • staff usually already know how to replace fans, filters, and cooling units;
  • the rack is easier to move or rearrange;
  • a redundant cooling unit can support several racks;
  • at moderate density, spare parts and contractors are easier to find.

The simplicity ends as airflow rises. Fan power, noise, duct size, and aisle area all increase. A raised floor that handled 8 kW per rack does not have to handle 40 kW after a server refresh. Teams often have to spread accelerators across several racks, leave rack units empty, or install rear-door heat exchangers. At that point, compare the cost of extra floor area and air-system modifications honestly with the cost of a liquid circuit.

A hot day tests more than the chiller. The design must withstand a dirty filter, the failure of one cooling unit, an aisle door opened for maintenance, and the maximum allowable IT load together in the agreed failure combination. If an N+1 failure pushes inlet temperature beyond the manufacturer's limit before controlled load reduction finishes, the redundancy exists only on the drawing.

Liquid removes density but does not eliminate air

Direct liquid cooling moves most heat from cold plates to a CDU and then to the building system, sharply reducing the airflow required through the rack. It does not automatically cool every component. Most current rack systems use a hybrid design: GPUs and CPUs reject heat to liquid, while fans cool memory, storage, power supplies, cables, and some network equipment.

Commercial descriptions often blur this distinction. A "liquid-cooled rack" can mean cold plates, a water-cooled rear door, immersion, or simply a CDU next to the rack. These options place different demands on servers, coolant, piping, and service. The technical specification should state the heat-transfer method and responsibility boundary rather than using the generic word "liquid."

A cold-plate system normally has two circuits. The Facility Water System removes heat at building level. The Technology Cooling System circulates treated coolant through the servers. A CDU separates them with a heat exchanger, maintains flow and pressure, filters the coolant, measures conditions, and controls pumps. Separation does not guarantee that leaks cannot happen, but it keeps building water out of the accelerators' fine channels.

The Open Compute Project's Cold Plate Cooling Loop Requirements calls for compatibility between the coolant and every wetted material, plus regular checks of water chemistry even when the original materials are compatible. That is more useful than the instruction to "fill it with distilled water." Chemistry changes because of oxygen, corrosion products, contamination during service, and mixtures of incompatible materials.

ASHRAE's new water classes put the upper supply-temperature limit in the name: W17, W27, W32, W40, W45, and W+; the lower limit for all these classes is 2°C. Warmer supply expands compressor-free operating hours and may make heat reuse easier. The equipment manufacturer still selects the class. You cannot send 40°C water to a server merely because the outdoor dry cooler works more efficiently there.

The density of current systems shows why this debate can no longer rest on the operations team's preference. NVIDIA documents the DGX GB200 NVL72 as a hybrid-cooled rack using about 120 kW: processors and accelerators connect to liquid manifolds, while air cools the remaining components. Liquid is part of the server system's architecture. Replacing it with powerful fans changes the certified design and is not a project option.

At lower density, a rear-door heat exchanger can be an intermediate option. The servers remain air cooled, while the door transfers heat from their exhaust into a water circuit. It reduces the room load but does not reduce airflow through the servers, and it adds weight, hoses, and rear-clearance requirements. Test it as a separate product for allowable airflow, water temperature, pressure drop, and behavior while open.

Compare cost over the same horizon

Air or liquid by calculation
GSE selects data center infrastructure for the equipment, room, and expansion plan.
Choose a solution

A sound comparison covers the entire site at the same load, redundancy level, and period instead of comparing the price of a cooling unit with the price of a CDU. Air cooling is often cheaper for one moderate rack in a prepared room. Liquid can win at high density, where floor area is expensive, over a long service life, or when the server architecture requires it.

Use the same formula for both options:

TCO = CAPEX + ENERGY + WATER + SERVICE + SPARES + DOWNTIME_RISK + REFIT - RESIDUAL_VALUE

For air, CAPEX includes cooling units, a chiller or DX section, aisle containment, a raised floor or ducts, electrical work, and controls. For liquid, add the CDU, manifolds, piping, heat exchanger, dry cooler or chiller, coolant treatment, leak sensors, drainage, and construction work. Both options need redundancy, commissioning, and monitoring.

Calculate ENERGY with an hourly model or at least several combinations of outdoor temperature and IT load. A nameplate PUE figure without measurement boundaries says little. Count compressors, server and cooling-unit fans, pumps, humidification, heating, and auxiliary equipment separately. Liquid often reduces fan energy, but excessive hydraulic resistance or redundant pumps that run continuously can consume part of the saving.

Do not equate WATER with liquid cooling. A closed cold-plate loop contains coolant but consumes almost none during normal operation. Water use usually comes from the external heat-rejection method, such as a cooling tower or adiabatic section. An air-cooled server connected to a water-cooled chiller and cooling tower can use more site water than a liquid-cooled server with a dry cooler. Compare the water balance of the whole system.

SERVICE includes coolant analysis, filters, pump seals, fans, heat-exchanger cleaning, sensor calibration, and mandatory routine tests. Under SPARES, enter a list rather than an arbitrary percentage: pump module, CDU controller, filter elements, quick disconnects, sensors, fans, and boards that can actually be replaced on site. Delivery time to Kazakhstan and customs logistics can matter more than a discount in the quotation.

Do not turn downtime risk into a fictitious exact figure. First describe each scenario, permitted duration, detection method, and impact on the compute queue. The business owner can then assign an hourly cost or maximum recovery time. This makes it clear whether a second CDU is worth the cost and whether a pump must be replaceable without shutting down the rack.

Run the calculation for at least three points: current utilization, the first phase's maximum configuration, and expansion after two or three accelerator generations. If air works today but requires a complete room rebuild for the next purchase, the low first-phase CAPEX may simply be a deferred invoice.

A weak response turns a leak into an incident

You cannot reduce liquid risk to zero, but you can divide it into prevention, early detection, automatic isolation, and safe recovery. A supplier's statement that "quick disconnects do not leak" is not a plan. A coupling can pass factory tests and later receive side load from a badly routed hose, collect dirt during service, or suffer seal damage after several connection cycles.

Prevention starts with material compatibility, permitted pressure, torque control, hose restraint, and a clear connection procedure. Place quick disconnects where a technician can see and hold both sides and where cables will not pull them down. Beneath the rack and CDU, provide a path to a sensor or tray so liquid cannot disappear unnoticed under neighboring equipment.

OCP's cold-plate requirements distinguish indirect detection through pressure drop or flow change from direct sensors that react to liquid. The document advises locating spot sensors or sensing cable where liquid will actually collect. This qualification matters: a sensor uphill on the floor or beyond the opposite edge creates confidence without reducing detection time.

An expensive rack merits detection at several levels:

  1. the node controller watches a local sensor and component temperatures;
  2. the rack monitors sensing cable, pressure, and manifold flow;
  3. the CDU compares supply and return and controls valves and pumps;
  4. the BMS receives an independent signal and starts the agreed response;
  5. the duty team sees the location, severity, and required action instead of a generic alarm.

The automatic response depends on the size of the leak. For a small signal, the system may isolate a branch and reduce node power. For a rapid pressure drop, it stops the pump, closes valves, and starts a controlled shutdown of affected equipment. Do not cut power to the entire room from one wet sensor without considering the consequences, but do not make an email to the duty technician the only response either.

Commissioning does not require pouring coolant onto a server. Test a sensor with the manufacturer's test liquid or method, simulate a flow change on a test circuit, interrupt the BMS connection under control, and confirm valve closure by position and actual flow. Record event-to-alarm time, isolation time, residual volume, IT-load response, and the return-to-service sequence.

Check insurance conditions and site rules before delivery as well. A room above an archive, clinical area, or electrical switch room may require double-contained pipe, trays, separate drainage, or a different location entirely. Leak risk depends on both probability and where the liquid will go.

The room changes in different ways

An expansion plan without dead ends
The integration project accounts for the next accelerator configuration and data center resources.
Choose a solution

Air cooling needs room for a large flow, while liquid cooling needs space and structural capacity for hydraulics. In either case, begin with a building survey: incoming electrical capacity, protection selectivity, floor loading, delivery route, door height, fire compartments, access to an outdoor plant area, and the ability to stop part of the system for repairs.

An air-cooled rack needs these room elements:

  • contained hot and cold air paths with blanking panels in empty rack units;
  • sufficient duct, grille, or raised-floor area;
  • outdoor-air filtration and filter pressure-drop monitoring;
  • rack inlet temperature sensors at the bottom, middle, and top;
  • a hot-air return without pockets or short-circuit paths.

Direct liquid cooling needs supply and return routes, space for the CDU, service clearances, fill and drain points, leak monitoring, drainage, and access to the heat exchanger. Compare the mass of the filled rack, manifold, and door with the floor's permitted point load. Check aisle width with doors open and hoses connected, not on an empty floor plan.

The circuit boundary should match the organizational boundary. Facilities staff are responsible for FWS temperature, pressure, and availability up to the agreed point. The IT team or supplier is responsible for the TCS, CDU, manifold, and servers. At the boundary, specify temperature range, flow, differential pressure, water quality, connector type, redundancy, and telemetry format. Otherwise, each team will point to the other team's gauge when equipment overheats.

Condensation is often a greater cold-plate hazard than a large leak because it starts quietly on a cold surface. Supply temperature must stay above room dew point by the agreed margin in every operating mode, including a door opened in summer, wet cleaning, and sensor failure. If the design requires water below dew point, it needs vapor barriers, insulation, and a separate analysis rather than faith in a dry climate.

The outdoor part of a liquid system needs freeze protection in winter. An air system with a direct economizer needs protection from intake air that is too cold or too dry. Both need transition-mode controls. The most dangerous weather is often not the design frost or heat, but a rapid change that leaves dampers, valves, and setpoints out of sequence.

Fire controls, emergency power-off, UPS systems, and the generator must account for the new thermal inertia. After utility power fails, accelerators continue to release heat, while pumps and fans may stop before jobs shut down cleanly. Calculate coast-down time, power for controllers and pumps, the power-capping sequence, and the ability to remove residual heat on generator power.

Maintenance determines availability more than technology

Density without a blind rebuild
The integrator checks power, heat rejection, and service boundaries before selecting the rack configuration.
Choose a solution

The better design is the one the local team can diagnose and restore in any season. Air equipment is more familiar, but it still requires maintenance. Liquid equipment adds new tasks, yet a well-designed system can locate failures more precisely and provide replaceable pump modules.

For air, operations checks filter pressure drop, airflow, tile and damper positions, containment seals, temperature maps, and the redundant cooling unit monthly or by condition. Check the path again after new cables are installed. One cable bundle in a server exhaust can spoil the result of a good calculation.

For liquid, the schedule includes coolant sampling, conductivity or other manufacturer-specified measurements, inhibitor condition, filter pressure drop, pump operation, coupling tightness, and sensor tests. Do not invent one chemistry specification for every system: water, water-glycol mixtures, and dielectric coolants need different methods, and the supplier defines warranty limits.

Write procedures for four routine jobs: connecting a new node, isolating a branch, replacing a pump module, and tracing a leak alarm. Each procedure states who authorizes the action, how residual pressure is measured, where coolant is drained, how connectors are protected from dirt, and which readings confirm a return to normal. Technicians should practice these operations on a training section before working over an accelerator rack.

Choose spares by recovery time. An air system normally keeps filters, fans, damper actuators, and essential control boards on site. A liquid system adds approved seals and quick disconnects, CDU filters, sensors, a pump module, and prepared spare coolant. A random hose from an industrial shop must never enter the server circuit.

Monitoring must place IT and facility measurements on one timeline. GPU temperature without supply temperature is of little use; a CDU alarm without rack power is equally incomplete. Retain supply and return temperatures, flow, pressure, valve position, pump status, dew point, inlet-air temperatures, rack power, fan speed, and throttling events. Incident review can then follow a causal chain.

Round-the-clock support matters only with an escalation matrix and real spare parts. GSE designs AI and data center infrastructure, manufactures servers in Kazakhstan, and operates a nationwide service network, so an integrated rack project can connect equipment selection with local service capability. The customer must still approve procedures, train the shift team, and run emergency drills.

Make the decision with thresholds and verify under load

Choose air when the servers are certified for air cooling, the verified airflow path has reserve, the room can survive an N+1 failure, and expansion will not force an immediate rebuild. Choose liquid when the equipment architecture requires it, the airflow physically cannot pass, density makes space and fan energy too expensive, or warm water provides a justified economizer benefit.

A hybrid design is often the exact answer rather than a compromise. Cold plates remove CPU and GPU heat, while a smaller air circuit serves the other components. An existing room can move in stages: one CDU and one liquid branch can start beside the retained air reserve. The stage still needs a final architecture, or temporary hoses and portable cooling units will remain for years.

Record the decision in a table with measurable thresholds:

CriterionAirDirect liquid
Maximum rack powerpasses the verified air balancepasses CDU and FWS thermal capacity
Design hot hourserver inlets remain within manufacturer limits under N+1supply temperature and residual air remain within limits under N+1
Dry and dusty periodfilters and humidification maintain the permitted rangethe air portion is protected and coolant stays within specification
Winter modeeconomizer does not overcool or overdry the roomoutdoor circuit is protected against freezing
Serviceteam replaces air components within the target timeteam isolates a branch and services the CDU within the target time
Expansionphysical airflow and floor-space reserve existpiping, CDU, heat-rejection, and electrical reserve exist

Hold two reviews before purchase. At the design review, the supplier presents inputs, heat and flow calculations, the control diagram, failure list, responsibility boundaries, and coolant specification. At the operational review, the operations team checks access, consumables, spares, instructions, telemetry, and contracted response time.

Acceptance must run at the design IT load, not with an idle rack. Tests include steady maximum load, transfer to a redundant pump or cooling unit, a filter at its permitted dirty limit, loss of one sensor, a simulated leak, changeover of the outdoor mode, and recovery after power loss. Define allowable temperature, response time, permitted power reduction, and the stop condition before each test.

If both options pass the thresholds, choose the one with lower risk in the ten-year cost model and local operation. If air fails the flow calculation, the debate about its lower cost is over. If liquid lacks approved chemistry, sensors, isolation, and a trained shift team, high density does not make it ready for service.

FAQ

At what rack power should I switch to liquid cooling?

There is no universal kilowatt threshold. You need to switch when the verified air path cannot provide the required flow at the permitted temperature under an N+1 failure, or when the rack manufacturer explicitly requires a liquid circuit.

Can a GPU rack use outdoor air for cooling in winter?

An air-side economizer can do this if controls manage temperature, dew point, filtration, and the rate of environmental change. Direct cold intake air without mixing can overcool or overdry server inlets.

Does Kazakhstan's dry climate help liquid cooling?

Yes. Dry air expands the operating range of dry and adiabatic coolers, but the benefit depends on the server's permitted supply temperature. It does not remove the need for dust protection, freeze protection, or dew-point control.

How dangerous is a leak in an accelerator rack?

The risk depends on volume, location, and response time. Compatible materials, correctly installed couplings, sensors at collection points, automatic valves, and a practiced isolation procedure greatly reduce the consequences.

Does direct liquid cooling require a chiller?

Not always. If the equipment permits warm enough supply water, a dry cooler or water-side economizer can remove heat without compressors for a significant period, but an hourly climate calculation must confirm it.

Does liquid cooling consume a lot of water?

A closed server loop consumes almost no coolant in normal operation. Most water use comes from an external cooling tower or adiabatic section, so compare the whole heat-rejection system rather than the rack alone.

Can I install a liquid-cooled rack in an existing data center?

Yes, if a survey confirms electrical capacity, floor loading, CDU space, pipe routes, drainage, heat rejection, and service clearance. The building often imposes the limit rather than the rack.

Which costs less to maintain, air or liquid cooling?

Air is usually simpler at moderate density in a prepared room. At high density, liquid may cut fan and floor-space costs, but it adds coolant analysis, pumps, filters, and sensors.

Can liquid cooling eliminate room air cooling completely?

Usually not. Many systems cool CPUs and GPUs with liquid while memory, storage, power supplies, and network components remain air cooled, so the design must include the residual air load.

How should the cooling system be accepted after installation?

Test it at the design IT load and reproduce agreed failures involving a pump, cooling unit, sensor, BMS link, and leak signal. The protocol should record temperatures, detection time, automatic response, and recovery conditions.