Skip to main content

Insights · 9 min

What Takes a GPU Hall Offline When Structured Cabling Is Wrong

The switches can be right. The GPUs can be on the floor. If the structured cabling and passive infrastructure are amateur, the cluster is inventory.

NP-CBL / Fabric planes

East-West · North-South · Management. Cabling builds the plants.

DWG-CBL-01 / NOT FOUR EQUAL JOBSSTRUCTURED CABLING BUILDS THESE PLANTSENTERPRISE / USERS / APPLICATIONSAPI GATEWAY / LB · FIREWALLS · CORENORTH-SOUTH / FRONT-END FABRICFRONT-END / APP / ENTERPRISEAI PLATFORM / DGX ENVIRONMENTNOT THE HALL · THE ACCELERATOR CELLEAST-WEST COMPUTE · QUANTUM-X / SPECTRUM-XGPU-TO-GPUDGX-TO-DGX · RDMASTORAGE / DATA · DEDICATED · COMPUTE ↔ HPSMGMT / OOB · NOT N-SDGXGPU rail 1DGXGPU rail 2DGXGPU rail 3DGXGPU rail 4OVERHEAD BASKETS / HOW THE PLANTS ARE BUILT / NOT A FIFTH FABRICEW · TRAY A · COMPUTE MPO · RAIL-ALIGNEDEW · TRAY B · STORAGE / DATA · OFF THE COMPUTE RAILSNS · TRAY C · FRONT-ENDMGMT · TRAY D · OOB / BMCN-S IS USERS AND APPSE-W IS GPU / DGX SCALE-OUTUNDERFLOOR = POWER / LIQUID · OVERHEAD = FIBER · SLACK AT RACK AND FRAMESTORAGE AND OOB INDEPENDENT

DWG-CBL-01 · laying out

EW / East-West

Hall-internal scale-out. Compute fabric first. Storage/data fabric on its own plane when dedicated. Not user traffic.

01 / East-West Compute Fabric

MPO/MTP trunks, 400/800G optics, rail-aligned

GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.

02 / Storage / Data Fabric

Separate MPO or LC plant onto the storage leaf when dedicated

Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails.

NS / North-South

Users, applications, API gateway/load balancer, enterprise/core, firewalls. Application-facing Ethernet. Not another GPU rail.

03 / North-South / Front-End Fabric

LC or lower-count MPO, 100G-class typical

Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O.

MGMT / Management

BMC, management, operational control. Never classify as application North-South.

04 / Management / OOB Fabric

Cat6A or dedicated 1/10G optics. Own patch field.

BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark.

NP-CBL-FAIL / Amateur plant

How a GPU hall stays dark

WRONG / ONE TRAYGPUSWITCH×Kink · mixed polarity · no slack · dirty MPOSPECIFIED / BY PLANEGPUEW COMPUTEGPU-TO-GPUDATA FABRICDEDICATEDN-S FRONTUSERS / APPSMGMT / OOBNOT N-SSlack loops · labeled · IL / polarity on record

Bend radius

A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes.

Dirty MPO

One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected.

Mixed polarity

Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks.

No slack

You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage.

Blocked airflow

DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem.

Untestable plant

No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days.

No as-builts

The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name.

Compute and storage on one tray

A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not.

People specify GPUs. They specify switches. They treat structured cabling as a fit-out. That is how a hall stays dark after the hardware has arrived.

This is not a branding problem. It is bend radius, dirty MPO, mixed polarity, no slack, blocked airflow, an untestable plant, no as-builts, and compute trunks sharing a tray with storage. Any one of those can zero a rail. Together they zero a hall.

Bend radius first. An MPO trunk kinked behind a PDU or a CDU manifold will pass a casual look. Under heat and vibration the insertion loss moves. You get CRCs and flaps. The job dies at forty minutes and everyone stares at NCCL.

Dirty MPO is the 400/800G classic. One contaminated lane on an MPO-12 or MPO-16 takes the port down. The switch looks guilty. The end-face was never inspected under IEC rules. A click is not a test.

Polarity is a drawing. Type A and Type B are not interchangeable. Mix a cassette and a trunk and some lanes light. Some stay dark. Firmware updates do not fix a crossed key.

No slack means the first service event is an outage. You cannot slide a GPU sled, reseat a NIC, or swap a QSFP without tensioning fiber. Strain relief belongs at the designed loop, not at the ferrule.

Airflow is a cable problem on air-cooled nodes. DAC and AEC are stiff and thick. Pack them into the rear chimney and inlet temperature rises. GPUs throttle, then shut down. Liquid-cooled NVL racks still need a path that does not fight hoses and manifolds.

InfiniBand vs Ethernet does not save you if the plants are mixed. Same cage class is not the same fabric. Rail-aligned layout means each GPU rail lands on the matching leaf. A 'nearest free port' install is how collectives look like software.

Storage/data, North-South front-end, and management/OOB are separate cable plants from the East-West compute fabric. NVIDIA SuperPOD reference architectures already split those switch planes. The trays have to split too. Pull a storage jumper off a shared basket and you can take training down with it. Kill OOB inside a compute bundle and you cannot even see the BMC. OOB is not application North-South.

Campus reach is a fiber number. NVIDIA's public DSX facilities overview puts a 500 m optical reach limit on the cluster interconnect spine. Hall placement and SMF plant have to land inside that class. A 700 m run of leftover MMF is not a CIN.

Test is the gate. End-face inspection. Insertion loss per lane, not a trunk average. Polarity check before optics. OTDR on SMF backbone. As-builts tied to the same IDs as the labels. If you cannot isolate a rail at 03:00 from the drawing, you do not have a plant. You have a pile of patch cords.

We have watched an unnamed operator buy a large CPU and GPU fleet, hand structured cabling to another firm, and sit offline for a long time. NeuronPlant did not pull that plant. We will not dress it up as a delivery. Recovery is a specified recable. It is not a software workaround.

NeuronPlant's position is simple. Specify the structured cabling and passive infrastructure with the fabric. Not after the GPUs land.