Skip to main content

Cautionary record · Operator withheld

High-density GPU hall / structured cabling failure

Challenge

The operator bought a large CPU and GPU fleet. Structured cabling was let to another firm as a fit-out. The hall has been offline for a long time. NeuronPlant was not the cabling contractor.

Scale

  • High-density GPU hall
  • Large CPU + GPU buy
  • 400/800G class fabric
  • Third-party structured cabling

If NeuronPlant is brought in

  • Not the original cabling contractor
  • Plant audit if engaged
  • Polarity / IL / OTDR map
  • As-built reconstruction
  • Per-fabric recable sequence
  • Test gates before optics reseat

What went wrong

  • MPO end-faces never inspected. Dirty lanes dropped 800G ports. Teams chased switch firmware.
  • Mixed polarity. Type A cassettes on Type B trunks. Half the rails stayed dark.
  • No slack. First sled pull tensioned fiber and took a rail down.
  • DAC bulk in the rear chimney. Air-cooled nodes throttled, then shut down.
  • Compute and storage trunks on one tray. A storage change yanked a compute plant.
  • No labels, no IL records, no as-builts. The plant cannot be isolated or retested.

NP-CBL-FAIL / Amateur plant

How a GPU hall stays dark

WRONG / ONE TRAYGPUSWITCH×Kink · mixed polarity · no slack · dirty MPOSPECIFIED / BY PLANEGPUEW COMPUTEGPU-TO-GPUDATA FABRICDEDICATEDN-S FRONTUSERS / APPSMGMT / OOBNOT N-SSlack loops · labeled · IL / polarity on record

Bend radius

A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes.

Dirty MPO

One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected.

Mixed polarity

Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks.

No slack

You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage.

Blocked airflow

DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem.

Untestable plant

No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days.

No as-builts

The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name.

Compute and storage on one tray

A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not.

Failure path

  1. 01Compute buy complete
  2. 02Amateur structured cabling
  3. 03Dirty MPO / mixed polarity
  4. 04Untestable trays
  5. 05Hall dark
  6. 06Recable is the recovery

Status

Still a cautionary record, not a NeuronPlant win. Compute sat while the original plant remained the plant of record. Recovery is a specified recable with test records. It is not a firmware patch.

NP / Intake

Tell us what you need to achieve with AI.

One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.

Start a Project →