# NeuronPlant / Plant notes Short technical notes on NVL72, power/cooling, fabrics, structured cabling, and steered RFPs. Canonical HTML: https://neuronplant.com/insights Collection file: https://neuronplant.com/llms/insight.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # What Does It Actually Take to Deploy an NVL72? It is not a server install. It is a power, water, floor, and fabric problem that happens to include 72 GPUs. An NVL72 rack is a liquid-cooled NVLink domain. Treat it as plant equipment, not as a dense 42U server. The first constraint is almost never the purchase order. It is facility water, rack power in the 100 kW class, floor loading, and a network design that does not starve the NVLink domain the moment training starts. Power: plan for a dedicated high-density feed the hall can present, A/B where the architecture requires it, and a PDU/busway story that can be maintained without taking the pod down. Nameplate and sustained draw are different numbers. Design to the one that shows up on the bill and the breaker. Managed vs unmanaged is a constraint: dual-cord training racks usually do not want outlet switching. Cooling: direct-to-chip liquid, CDU, and a heat-rejection path back to the facility water or dry coolers the hall supplies. If the hall cannot present the water, the rack does not ship, or it ships and sits. Redundancy lives in the loop, the CDU, and the rejection plant, not in a spare fan. Network: the scale-up domain is NVLink. The scale-out domain is still a fabric you have to design: 800G class East-West compute, with storage/data and management/OOB as separate planes, and North-South for users and apps. Cabling, optics, and cable management are engineering, not day-two tidy-up. Operations: you need a runbook before first power-on. Leak detection, isolation valves, firmware, BMC, and a commissioning sequence that proves cooling before it proves FLOPS. NeuronPlant's position is simple. Specify the rack only after the room we occupy, the loop, and the feed can take it. We engineer the cell to fit. We do not become the hall. Canonical: https://neuronplant.com/insights/deploy-nvl72 # Power and Cooling Requirements for Modern AI Factories Accelerator density moved AI out of the IT closet and into industrial infrastructure. GPU is the common case. The planning units are kilowatts, liters, and months. Traditional IT racks lived in the 5-15 kW range. AI racks did not stay there. Air-cooled high-density GPU systems already push past what a standard raised floor and CRAH fleet were designed to do. Liquid is not a preference at that point. It is the load path. Power planning starts at what the hall can present, not at the PDU SKU. Available kW, interconnection lead time the hall owner carries, transformer capacity, UPS topology, and the decision between centralized and rack-level redundancy determine the cell more than the SKU list. At the rack, a PDU is how power is distributed. Managed means a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Specify that job when you need to see kW, phase balance, and headroom versus nameplate (training loads are not nameplate), when the hall is multi-tenant or billed, or when the board and BMS do not already meter at that grain. Managed is not automatic. Dual-cord GPU and compute racks do not want outlet switching. You do not power-cycle a training node from a PDU web UI. If EPMS or BMS already meters the feed, a metered PDU can be redundant. The PDU management network is another plane to design, secure, and cable. Unmanaged, or metered but not switched, is a valid spec. A useful rule: if you cannot draw the path the hall presents from its feed to the compute rack, including maintenance bypass, you do not yet have an architecture. Redundancy that exists only on a slide will fail the first time a feed is isolated. Drawing that path is not NeuronPlant building a substation. Cooling planning is a heat-rejection problem. CDUs, facility water temperatures, approach, and what happens when a loop is isolated matter as much as cold-plate selection. Design the failure, then design the steady state. PUE still matters, but for AI factories the binding constraints are often time-to-power and time-to-water: clocks that belong to the hall owner or the custom-system supplier. A perfect cooling plant that arrives a year after the accelerators is not an AI factory. It is inventory. The engineering sequence we use is feasibility first: power, water, space, floor, and connectivity the hall can present. Architecture of the accelerator cell second. Procurement third. That order is what keeps a multi-million-dollar bill of materials from arriving at a room that cannot run it. NeuronPlant does not pour the room to rescue a bad sequence. Canonical: https://neuronplant.com/insights/power-cooling-ai-factories # Ethernet vs InfiniBand for AI Infrastructure The fabric is part of the model runtime. Choose it for collectives, storage, and operations, not for a vendor slogan. Training workloads are sensitive to latency and to congestion during collective operations. Inference and RAG are often more sensitive to isolation, East-West fairness, and the path to storage. One fabric rarely serves every plane well if it is treated as a single flat network. InfiniBand remains a proven choice for tightly coupled training domains. It is not automatically the right choice for every AI factory. Ethernet at 400/800G, with a lossless or properly engineered congestion story, is now a serious training fabric, and often a better fit where the operator already runs Ethernet at scale. NVIDIA SuperPOD reference architectures name Quantum-X800 Q3400 or Quantum-2 QM9700 for compute, QM9700 or Spectrum SN5600 for storage, and SN5610 for user/NFS paths. What actually decides the design: message size and parallelism strategy, number of endpoints, storage protocol, in-house operational skill, and whether the same fabric is expected to carry tenant traffic. We separate planes. East-West compute, storage/data when dedicated, North-South front-end, and management/OOB are different networks even when they share a vendor. Collapsing them to save ports is how you create an outage that looks like a 'model problem'. Management/OOB is never application North-South. There is no poster-child answer. There is a workload, a scale, and an operations team. The architecture should be able to say why InfiniBand, why Ethernet, or why both, in writing, with a diagram, before anyone orders optics. Canonical: https://neuronplant.com/insights/ethernet-vs-infiniband # What Takes a GPU Hall Offline When Structured Cabling Is Wrong The switches can be right. The GPUs can be on the floor. If the structured cabling and passive infrastructure are amateur, the cluster is inventory. People specify GPUs. They specify switches. They treat structured cabling as a fit-out. That is how a hall stays dark after the hardware has arrived. This is not a branding problem. It is bend radius, dirty MPO, mixed polarity, no slack, blocked airflow, an untestable plant, no as-builts, and compute trunks sharing a tray with storage. Any one of those can zero a rail. Together they zero a hall. Bend radius first. An MPO trunk kinked behind a PDU or a CDU manifold will pass a casual look. Under heat and vibration the insertion loss moves. You get CRCs and flaps. The job dies at forty minutes and everyone stares at NCCL. Dirty MPO is the 400/800G classic. One contaminated lane on an MPO-12 or MPO-16 takes the port down. The switch looks guilty. The end-face was never inspected under IEC rules. A click is not a test. Polarity is a drawing. Type A and Type B are not interchangeable. Mix a cassette and a trunk and some lanes light. Some stay dark. Firmware updates do not fix a crossed key. No slack means the first service event is an outage. You cannot slide a GPU sled, reseat a NIC, or swap a QSFP without tensioning fiber. Strain relief belongs at the designed loop, not at the ferrule. Airflow is a cable problem on air-cooled nodes. DAC and AEC are stiff and thick. Pack them into the rear chimney and inlet temperature rises. GPUs throttle, then shut down. Liquid-cooled NVL racks still need a path that does not fight hoses and manifolds. InfiniBand vs Ethernet does not save you if the plants are mixed. Same cage class is not the same fabric. Rail-aligned layout means each GPU rail lands on the matching leaf. A 'nearest free port' install is how collectives look like software. Storage/data, North-South front-end, and management/OOB are separate cable plants from the East-West compute fabric. NVIDIA SuperPOD reference architectures already split those switch planes. The trays have to split too. Pull a storage jumper off a shared basket and you can take training down with it. Kill OOB inside a compute bundle and you cannot even see the BMC. OOB is not application North-South. Campus reach is a fiber number. NVIDIA's public DSX facilities overview puts a 500 m optical reach limit on the cluster interconnect spine. Hall placement and SMF plant have to land inside that class. A 700 m run of leftover MMF is not a CIN. Test is the gate. End-face inspection. Insertion loss per lane, not a trunk average. Polarity check before optics. OTDR on SMF backbone. As-builts tied to the same IDs as the labels. If you cannot isolate a rail at 03:00 from the drawing, you do not have a plant. You have a pile of patch cords. We have watched an unnamed operator buy a large CPU and GPU fleet, hand structured cabling to another firm, and sit offline for a long time. NeuronPlant did not pull that plant. We will not dress it up as a delivery. Recovery is a specified recable. It is not a software workaround. NeuronPlant's position is simple. Specify the structured cabling and passive infrastructure with the fabric. Not after the GPUs land. Canonical: https://neuronplant.com/insights/passive-plant-offline # A Steered RFP Is Not a Plant Spec A factory tender already aimed at one vendor is a capture. Write the constraints first. Leave the SKU until the plant can take it. An RFP that already names the winner is not a specification. It is a capture dressed as process. The first constraint is the plant: site, power, cooling, compute, network, structured cabling, storage, platform, operate. If those are not written, the SKU list is a shopping cart. Advisors get paid. Some get paid more when a particular BOM ships. That is not a secret. What fails is hiding that money path from the customer. NeuronPlant's position is simple. The customer sees the same constraints, options, and money path we see. We do not steer a plant toward the SKU that pays the advisor. Integrity before margin on a steered BOM. If the two conflict, we lose the margin. We do not lose the spec. No games. No tricks. Transparency is the commercial method, not a footnote. Canonical: https://neuronplant.com/insights/steered-rfp