14 min read

Data Center Liquid Cooling Operations

A day to day operating guide for liquid cooled data center halls, covering CDU and coolant loop monitoring, service boundaries, leak response, and preventive maintenance.

ByAndré Ribeiro· Founder, Obelinf
Data Center Liquid Cooling Operations
Data Center Liquid Cooling Operations · September 27, 2026
On this page

Liquid cooling makes a dense rack possible, but it also adds a fluid system to the operating environment. A stable hall depends on more than a CDU that is powered on. Operators need to know which team owns each pipe and coupling, what normal flow and pressure look like for each pod, how alarms reach a person who can act, and which valves can isolate a problem without creating another one.

This guide is about running a liquid cooled hall after commissioning. It follows the loop from facility water to the server, lays out practical monitoring and maintenance routines, and gives a response sequence for leak and cooling alarms. The exact setpoints and work steps must come from the equipment manufacturers and the site’s approved procedures, because coolant, CDU, and cold plate designs vary.

Map the loop and service boundaries

Most direct to chip installations have at least two fluid domains. The facility water system carries heat from the CDU toward the building’s heat rejection plant. The technology cooling system, sometimes called the secondary loop, carries conditioned coolant from the CDU to row or rack manifolds and then to server cold plates. A heat exchanger inside the CDU transfers heat between those circuits while keeping their fluids separate.

Facility water and technology cooling loops meet at the CDU, with service demarcations shown at each connection Two fluid domains, one heat transfer boundary Facility water plant, headers, valves CDU heat exchanger, pumps, controls Technology loop manifold, cold plates facility demarcation IT loop demarcation Write down the actual flange, valve, or quick connect that defines each team’s work boundary. Return flow carries heat back toward the CDU; arrows show supply and return paths.

The word “facility side” is not a work order boundary by itself. The lease, operating agreement, and commissioning documents should identify the physical demarcation: for example, the CDU inlet and outlet flanges, the rack manifold isolation valves, or the server quick disconnects. For each point, name who can operate the valve, who can repair it, who supplies replacement coolant, and who has authority to stop the affected IT load.

Keep a loop drawing that matches the installed valves and labels. Include CDU identifiers, branch and rack names, supply and return direction, isolation points, drains, vents, leak sensors, and the location of makeup fluid. If a contractor’s drawing disagrees with the as built labels, resolve that before the next service window rather than during an incident.

Turn commissioning into an operations handoff

Routine operations start with a deliberate handoff, not with the first shift after a contractor leaves. The facility operator, CDU supplier, IT equipment owner, and any cooling service provider should agree what “ready for service” means. The handoff package should include approved drawings, equipment schedules, control sequences, coolant specifications, commissioning results, alarm lists, warranty contacts, and the procedures for filling, venting, flushing, isolation, and draining. Operators should know which work is permitted in house and which work requires a qualified service technician.

Walk the installed system with the drawings in hand. Follow supply and return lines from the facility connection through the CDU and onward to each row or rack branch. Check that labels are readable from the normal approach, valves match the drawing, sensors have unique names, and drains or vents are accessible without moving energized equipment. Record any mismatch as a handoff issue with an owner and due date. A drawing that is correct in principle but impossible to reconcile with the installed labels will not help an overnight responder.

Review the commissioning evidence for both normal and abnormal operation. It should show that the loop was flushed and filled using the specified fluid, air was removed, flow was balanced, pressure integrity was checked, controls were configured, and monitoring points report plausible values. Review what happens on a pump failure, loss of facility water, high supply temperature, leak detection, communication loss, and power interruption. If redundancy is claimed, confirm the switchover behavior and the remaining capacity under the expected load. Do not assume that two pumps or two CDUs provide useful redundancy until the controls, valves, power feeds, and operating procedure support that result.

Capture the initial baseline at more than one operating condition when practical. A quiet commissioning load does not prove the loop is healthy near design load. Record readings during staged load increases and after stabilization, and note ambient conditions and any facility plant changes. When equipment is added later, repeat the relevant measurements and update the baseline. Preserve the dated snapshots, since gradual changes are easier to spot when the team can compare today’s condition with the original accepted condition and the last maintenance event.

Establish an operating baseline

At handover, capture normal readings at a known IT load and a known facility condition. Record CDU supply and return temperatures, loop flow, differential pressure, pump speed and state, filter differential pressure, fluid level or makeup volume, valve position, and active alarms. Note ambient and dew point conditions where condensation is possible. Preserve the CDU and server vendor limits beside the baseline so an operator can distinguish “different from normal” from “outside equipment limits.”

Do not copy a generic target from another hall. Required flow depends on the cold plates and the heat being removed. Pressure limits depend on the weakest rated part of the connected loop. Supply temperature must satisfy IT equipment requirements and remain appropriate for condensation control at the local surfaces. Use the qualified design and equipment documentation to set warning and action thresholds.

Use primary engineering references to check the design assumptions: ASHRAE’s data center handbook chapter describes liquid cooling configurations and facility interfaces, while the Open Compute Project CDU workstream publishes integration guidance. For operating limits and service procedures, follow the manual for the installed CDU and IT equipment, such as this Vertiv CDU operation and maintenance guide, rather than treating another model’s values as universal.

Trend readings together. A high return temperature alone may follow a change in workload, but a rising return temperature with falling flow and increasing pump speed points toward a restriction or an emerging hydraulic issue. Compare rack level measurements with CDU totals to identify a branch that is out of balance. Alert on both hard limits and meaningful deviation from the commissioned baseline, with enough persistence to avoid paging on a momentary sensor blip.

Interpret the monitoring signals together

The CDU is a useful measurement point, but it cannot describe every rack branch by itself. A facility water problem can affect several CDUs at once, while a closed branch valve or clogged strainer can affect only one row. Collect readings at the points that let you separate those cases: facility water entering and leaving the CDU, technology loop supply and return, CDU flow and differential pressure, pump command and feedback, branch flow where available, and rack or server thermal telemetry. Keep timestamps aligned so the team can see which readings changed first.

Build alerts around relationships as well as individual limits. For example, rising technology loop supply temperature with stable facility water may indicate a CDU heat exchanger or control issue. A rack branch with falling flow while the overall CDU flow remains steady may indicate a local restriction or valve position change. A rising pump speed that no longer restores expected flow can suggest increasing resistance, a sensor fault, or a pump problem. These patterns narrow the investigation; they do not prove a cause, so operators should verify them against the device display, the trend history, and the approved troubleshooting steps.

Watch the difference between a real value and a missing value. A flat trend can mean stable operation, but it can also mean a failed sensor, stale gateway, or lost network connection. Monitoring should expose last update time, sensor quality, communication state, and alarm suppression. If a point goes stale, route it as a telemetry fault with an owner and restoration target. A cooling system with blind sensors should not continue to be reported as fully monitored.

Keep process values, equipment alarms, and operator annotations together. When a technician changes a filter, adjusts a valve, tops up coolant, or switches a pump, record the event with a timestamp. Without that history, the next shift may interpret an expected post maintenance change as a new failure, or miss a slow deterioration because a setpoint adjustment concealed the trend. Trends become more valuable when each operational change can be placed beside the readings it affected.

Daily checks and alarm ownership

At the start of a shift, review the CDU state, active and recently cleared alarms, supply temperature, flow, pressure differential, pump status, filter status, and leak detection zones. Confirm the monitoring system has received current values rather than treating a stale last reading as healthy. Walk the accessible equipment and look for fluid at couplings, drains, valve stems, CDU pans, and manifold connections. Check for unusual pump sound, vibration, or a change in condensation on nearby surfaces.

Every alarm should have an owner, a first action, and an escalation path. The operator seeing a low flow alarm needs to know which dashboard to check, which team to call, and whether the site design permits a standby pump or another CDU to take the load. A sensor alarm should identify its zone and equipment label, not just say “liquid detected.” Test the path from sensor to notification and record the response, including who acknowledged it.

Separate information alerts from action alarms. A slow drift in filter differential pressure can create a maintenance work order; a confirmed leak sensor, loss of flow, or high temperature approaching an IT limit needs immediate incident handling. Do not let a dashboard summary hide the local CDU alarm or its event history.

Use a repeatable check cadence so that each shift does not invent its own definition of “looked at.” The interval below is a starting point for the operating plan, not a substitute for vendor instructions or a site risk assessment.

Cadence Checks and records Why it matters
Each shift Review alarms, stale points, supply and return temperatures, flow, pressure, pump state, leak zones, and recent operator notes. Walk accessible CDU and manifold areas. Finds acute changes and confirms the monitoring path is alive.
Weekly Review trend exceptions, recurring acknowledgements, filter differential pressure, makeup fluid, sensor faults, and open corrective work. Confirm leak zones have not been bypassed. Reveals repeated alarms and slow changes that a shift view can miss.
Monthly or per site plan Reconcile equipment labels and open work against the current loop drawing. Review maintenance due dates, spare parts, supplier advisories, and changes in IT load or coolant requirements. Keeps the operating model aligned with the installed system.
At approved test intervals Test leak detection, notification, standby equipment behavior, and documented valve or shutdown sequences using the site’s safe test method. Demonstrates that the response path works before an actual incident.

For every check, define where the source of truth lives and what evidence counts as completion. If a CDU’s own display reports a fault but the central system is green, the discrepancy should become a ticket, not an informal note. If a leak detection circuit is temporarily bypassed for service, log the affected zone, compensating measures, responsible person, and restoration deadline. An untracked bypass can quietly turn a monitored system into an unmonitored one.

Troubleshoot by symptom and scope

Start with the scope of the change. Did one rack, one branch, one CDU, or several CDUs change at the same time? A single affected rack points the first checks toward its branch valve, quick connections, local manifold, and workload. Several rows tied to one CDU point toward CDU operation or a shared header. A number of CDUs changing together points upstream toward facility water, controls, electrical supply, or a shared monitoring issue. This scoping step helps bring the right team in early without making the operator guess at the repair.

For low flow, compare the local branch reading with CDU total flow, pump command and feedback, differential pressure, and valve position. Check whether a recent service event or new rack connection lines up with the change. If the pump is commanded to a higher speed but flow is still low, stop treating it as a routine load fluctuation and follow the vendor troubleshooting sequence. Avoid raising pressure or repeatedly restarting pumps unless the approved controls and operating procedure explicitly call for it.

For rising temperatures, distinguish a change in heat load from a change in cooling delivery. Look at server workload and rack power alongside coolant supply temperature, return temperature, and flow. If the load rose and flow and supply conditions remain normal, the hall may be responding to a workload event. If temperature rises while flow falls, check the hydraulic path and CDU alarms. If facility water temperatures or availability changed at the same time, involve the building plant team. Use the IT equipment telemetry to understand remaining thermal headroom, but follow the site’s action thresholds before the servers protect themselves by throttling.

For pressure drift or repeated makeup fluid, compare the affected loop section with other sections and inspect the places where fluid can escape or air can enter. Repeatedly adding fluid without finding the reason can hide a developing leak, introduce contamination, and complicate the incident record. Treat a new makeup event as a logged exception. Verify the amount, fluid type, source container, loop segment, and approval, then check whether the pressure trend returns to its previous behavior.

For a single leak sensor alarm with no visible fluid, keep the event open until the correct zone has been inspected safely and the sensor is shown to be dry and functional. Condensation, a displaced sensing cable, a spill from unrelated work, or a failed sensor are all possible explanations. The response should confirm the input at the local controller, identify the physical sensor, and document why the alarm was cleared. Resetting from a remote panel without checking the area removes evidence and can leave an actual leak undetected.

Respond to leaks and cooling loss

Treat a leak alarm as genuine until the sensor location has been checked. A small coolant loss can become a larger discharge when a pressurized connection moves, and a wet floor or electrical hazard can make access unsafe. Follow site emergency procedures, keep people clear of hazards, and bring facilities, IT, and the designated liquid cooling service team onto the incident bridge.

Liquid cooling alarm response from detection through safe isolation and documented restart Respond by zone, then contain the affected segment Alarm arrivessensor and zone Make area safepeople and power risk Confirm conditionvisual check if safe Isolate per plansmallest safe segment Repair, provethen restore Never loosen a pressurized coupling or improvise a valve sequence. Only the approved runbook and trained responders define isolation, shutdown, and restart steps.

Use the site runbook to identify the sensor zone and the nearest approved isolation points. If the design includes automatic shutoff, verify what it actually closes and what cooling remains available to the unaffected racks. If manual isolation is required, use only the documented sequence and trained personnel. Do not disconnect a quick coupling, open a drain, or stop a pump based on guesswork: an incorrect action can release more fluid or remove cooling from a healthy branch.

For a leak, contain the source if the procedure permits, protect nearby electrical equipment, and preserve cooling for unaffected IT. For loss of flow or a rising temperature without visible fluid, compare pump state, valve position, pressure differential, facility water availability, and rack level flow. Escalate to controlled IT load reduction or shutdown at the thresholds defined by the server and site procedures. Do not wait for thermal throttling to become the alarm strategy.

Before restart, repair the failed part, inspect nearby connections, restore fluid only with the specified coolant, remove trapped air using the approved method, and verify pressure, flow, temperature, and leak sensors. Clear alarms only after the readings are stable and the affected equipment owners agree to return the segment to service. Record the cause, isolation point, fluid added, equipment affected, and follow up work.

Preventive maintenance without surprise outages

Build the maintenance calendar from the CDU, pump, server, sensor, and coolant supplier requirements. Check filters and strainers against differential pressure and the recommended service interval. Confirm leak sensors, alarms, isolation valves, and notification paths on a planned schedule. Inspect hoses, couplings, supports, insulation, drip trays, drains, and accessible pipe joints for wear, movement, corrosion, deposits, or signs of prior moisture.

Fluid sampling should be planned with the coolant supplier and equipment manufacturers. Track the specified chemistry and contamination indicators, such as conductivity or particle levels when the system requirements call for them. Keep sample locations, laboratory results, fill records, and any top up volume together. Do not mix coolant formulations or add treatment chemicals without compatibility approval; fluid quality can affect metals, elastomers, filters, cold plates, and warranty coverage.

Treat work on a CDU or live loop as controlled maintenance. Review the impact on redundancy and remaining cooling capacity, notify affected teams, approve the method statement, and use lockout and tagout wherever required. Confirm that drain, vent, and isolation points are identified before work starts. After work, inspect the joint, verify readings under load, test the relevant alarm path, and update the as built drawing and maintenance record.

Make the runbook usable at 3 AM

A useful operating record fits the response. Keep the current loop diagram, normal ranges, vendor limits, sensor map, valve sequence, coolant type, spill supplies, escalation contacts, and restart criteria where the on call operator can reach them. Use stable equipment names across the CDU display, monitoring dashboard, rack labels, and work orders. During a planned exercise, have someone unfamiliar with the loop locate a simulated alarm and explain which team owns the next action.

The daily discipline is straightforward: trend the whole loop, respect the service demarcations, investigate anomalies before they become IT alarms, and isolate only what the approved design allows. Liquid cooling operations become manageable when the facility, CDU, rack, and server boundaries are visible to the same people who respond to them.

Frequently Asked Questions

What should operators monitor on a data center CDU?
Trend supply and return temperatures, flow, differential pressure, pump state, filter status, alarms, and makeup fluid. Compare readings with commissioned baselines and the limits for the connected IT equipment, since there is no single safe setpoint for every CDU and cold plate combination.
Who is responsible for a liquid cooling leak in a data center?
Responsibility depends on the agreed service boundary. Facility teams commonly own facility water up to the CDU, while the CDU, technology loop, rack manifold, and server connections may belong to different facility, provider, or tenant teams. Document the demarcation points and escalation contacts before deployment.
What should you do when a liquid cooling leak alarm activates?
Treat the alarm as real, identify its exact sensor zone, protect people, and follow the approved incident runbook to isolate only the affected segment if it is safe to do so. Do not disconnect a pressurized coupling or restart equipment until the leak source is repaired, the area is dry, and the responsible teams clear the system.
How often should a data center coolant loop be maintained?
Use the CDU, server, and coolant supplier maintenance schedules, then add routine checks for trends, visible leaks, sensor health, and fluid condition. Sampling and filter service intervals depend on coolant chemistry, materials, contamination risk, and operating history, so set them from commissioning data and supplier guidance.

Stop reaching for a spreadsheet

Obelinf keeps every subnet, device, circuit, and rack in one live source of truth, with audit logs and a topology view. Free for personal use.

Related Articles