Skip to main content

Insights · 7 min

Power and Cooling Requirements for Modern AI Factories

Accelerator density moved AI out of the IT closet and into industrial infrastructure. GPU is the common case. The planning units are kilowatts, liters, and months.

Traditional IT racks lived in the 5-15 kW range. AI racks did not stay there. Air-cooled high-density GPU systems already push past what a standard raised floor and CRAH fleet were designed to do. Liquid is not a preference at that point. It is the load path.

Power planning starts at what the hall can present, not at the PDU SKU. Available kW, interconnection lead time the hall owner carries, transformer capacity, UPS topology, and the decision between centralized and rack-level redundancy determine the cell more than the SKU list.

At the rack, a PDU is how power is distributed. Managed means a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Specify that job when you need to see kW, phase balance, and headroom versus nameplate (training loads are not nameplate), when the hall is multi-tenant or billed, or when the board and BMS do not already meter at that grain.

Managed is not automatic. Dual-cord GPU and compute racks do not want outlet switching. You do not power-cycle a training node from a PDU web UI. If EPMS or BMS already meters the feed, a metered PDU can be redundant. The PDU management network is another plane to design, secure, and cable. Unmanaged, or metered but not switched, is a valid spec.

A useful rule: if you cannot draw the path the hall presents from its feed to the compute rack, including maintenance bypass, you do not yet have an architecture. Redundancy that exists only on a slide will fail the first time a feed is isolated. Drawing that path is not NeuronPlant building a substation.

Cooling planning is a heat-rejection problem. CDUs, facility water temperatures, approach, and what happens when a loop is isolated matter as much as cold-plate selection. Design the failure, then design the steady state.

PUE still matters, but for AI factories the binding constraints are often time-to-power and time-to-water: clocks that belong to the hall owner or the custom-system supplier. A perfect cooling plant that arrives a year after the accelerators is not an AI factory. It is inventory.

The engineering sequence we use is feasibility first: power, water, space, floor, and connectivity the hall can present. Architecture of the accelerator cell second. Procurement third. That order is what keeps a multi-million-dollar bill of materials from arriving at a room that cannot run it. NeuronPlant does not pour the room to rescue a bad sequence.