NVIDIA’s DGX H100 data center guide puts the maximum power demand of one system at 10.2 kW. Four systems bring the server portion of a rack to 40.8 kW before switches, storage and management hardware are added.
Those figures do not set a universal density for AI facilities. They show why planners should start with the final equipment list and peak demand instead of multiplying an older server-room average by the number of new racks.
Build a capacity record for every rack
List every server, switch, storage appliance, rack power distribution unit and cooling component planned for the cabinet. For each item, record quantity, rated and expected power, input-voltage range, plug type, power-supply count, weight, rack units, airflow direction, cooling connection and network ports.
Keep two totals: the credible peak IT load used to size the facility path, and the measured or estimated operating load used for energy and cost forecasts. Add the power drawn by rack-level cooling and network equipment, then reserve explicitly approved space and capacity for expansion rather than leaving an undefined margin.
NVIDIA’s planning guide lists a DGX H100 at a maximum of 10.2 kW and 130.45 kg, and shows rack configurations ranging from 10.2 kW for one system to 40.8 kW for four. The same guide says a power or cooling constraint can force equipment into more racks, which can change cable lengths and network performance.
Trace redundant power to the upstream source
Map each electrical path from the utility or on-site generation through transformers, uninterruptible power supplies (UPSs), switchgear, busways, floor distribution and rack power distribution units to each server power supply. Record the rated capacity, usable capacity, protection device and maintenance state at every stage.
Two plugs on a server do not prove end-to-end redundancy. Paths labeled A and B may converge at one UPS, switchboard or transformer, leaving a common failure point. The design review should show which loads transfer during maintenance or failure and whether the surviving equipment can carry them.
NVIDIA’s H100 electrical guide specifies three rack power paths for its N+1 design, requires each path to support half of expected peak rack demand and tells installers to verify redundancy by switching off the breakers feeding each rack power strip. That topology is specific to the cited system; other hardware must be checked against its own power-supply and support requirements.
Match cooling to the heat that remains in the room
For air-cooled equipment, calculate airflow and heat rejection at the proposed inlet temperature, altitude and rack density. Confirm aisle layout, blanking panels and containment, then test for recirculation and hot spots at the intended load.
A liquid-cooling plan needs supply and return temperatures, required flow, pressure limits, fluid chemistry, material compatibility, filtration and service clearances. It should also identify the coolant distribution unit, manifolds, couplings, isolation valves, leak sensors, containment and the response to pump or power loss.
The Open Compute Project’s cold-plate requirements say hybrid systems still need room air conditioning, while coolant distribution units manage pressure, flow, temperature, cleanliness and leak detection. The document also calls for a spill and leak plan covering detection, intervention, containment and pump failures.
Do not assume that adding cold plates removes every air-side load. Memory, power electronics, network devices and other components may still reject heat into the room unless the rack design captures it elsewhere.
Check weight, delivery routes and network geometry
Add the cabinet, compute hardware, power equipment, coolant, manifolds and cabling when calculating rack weight. A structural professional should assess both the installed load and the rolling load along the route from the loading area to the final position. Door widths, ramps, lifts, turning space and floor transitions belong in the delivery review.
Place compute, storage and switches as one system rather than treating the network as a later cable order. Document port counts, link speeds, oversubscription, cable types, maximum path lengths, management networks and the effect of losing a link or switch. Spreading machines across extra rows may solve a cooling limit while adding cable distance or switching layers.
Commission the facility in a fixed sequence
Freeze the bill of materials, firmware versions and rack positions before integrated testing. Record the expected result, acceptance threshold, responsible person and recovery procedure for every test.
- Verify equipment labels, power connections, cooling connections and network ports against the final drawings.
- Measure voltage, current and phase balance, then remove each claimed redundant power path in turn.
- Apply the planned load and record inlet temperatures, component temperatures, coolant flow, pressure and room hot spots.
- Test compute, storage and network throughput with the intended topology and software versions.
- Simulate the loss of one rack, one switch and each cooling or power component that the design claims to tolerate.
- Trigger power, temperature, flow, leak and network alarms and confirm that they reach the assigned operators.
- Save electrical single-line diagrams, piping diagrams, cable maps, test results, change records and recovery steps.
Problems found before delivery can change a purchase order or rack layout. Problems found after installation may require new electrical work, piping, network paths or structural review while expensive equipment remains idle.
Frequently asked questions
Should power capacity use the rated or average load?
Use a defensible peak figure from the equipment specification and intended configuration when checking capacity and failure states. Keep measured or estimated operating load separate for energy and cost forecasts.
Do dual power supplies provide full redundancy?
Not by themselves. Trace both feeds to their upstream sources and identify every point where they share a UPS, switchboard, transformer or generator.
Does liquid cooling eliminate room air conditioning?
Not for a typical hybrid system. The facility must still remove heat from components and equipment that do not transfer all of their heat to the liquid loop.
Can commissioning wait until all servers arrive?
Final integrated tests require the installed system or representative loads, but power, cooling, weight and delivery constraints should be resolved during procurement and facility design.
Sources and further reading
- Planning a Data Center Deployment(NVIDIA)
- Electrical Specifications(NVIDIA)
- Liquid Cooling Cold Plate Requirements Document(Open Compute Project)