Skip to main content

Insights · 8 min

What Does It Actually Take to Deploy an NVL72?

It is not a server install. It is a power, water, floor, and fabric problem that happens to include 72 GPUs.

An NVL72 rack is a liquid-cooled NVLink domain. Treat it as plant equipment, not as a dense 42U server.

The first constraint is almost never the purchase order. It is facility water, rack power in the 100 kW class, floor loading, and a network design that does not starve the NVLink domain the moment training starts.

Power: plan for a dedicated high-density feed the hall can present, A/B where the architecture requires it, and a PDU/busway story that can be maintained without taking the pod down. Nameplate and sustained draw are different numbers. Design to the one that shows up on the bill and the breaker. Managed vs unmanaged is a constraint: dual-cord training racks usually do not want outlet switching.

Cooling: direct-to-chip liquid, CDU, and a heat-rejection path back to the facility water or dry coolers the hall supplies. If the hall cannot present the water, the rack does not ship, or it ships and sits. Redundancy lives in the loop, the CDU, and the rejection plant, not in a spare fan.

Network: the scale-up domain is NVLink. The scale-out domain is still a fabric you have to design: 800G class East-West compute, with storage/data and management/OOB as separate planes, and North-South for users and apps. Cabling, optics, and cable management are engineering, not day-two tidy-up.

Operations: you need a runbook before first power-on. Leak detection, isolation valves, firmware, BMC, and a commissioning sequence that proves cooling before it proves FLOPS.

NeuronPlant's position is simple. Specify the rack only after the room we occupy, the loop, and the feed can take it. We engineer the cell to fit. We do not become the hall.