01 / East-West Compute Fabric
MPO/MTP trunks, 400/800G optics, rail-aligned
GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.
People specify GPUs and switches. They treat structured cabling as a fit-out. That is how a hall stays dark after the hardware has arrived.
Fiber (MMF/SMF, polarity, MTP/MPO, 400/800G optics, reach). DAC, AEC, AOC. InfiniBand vs Ethernet layouts. Trays for East-West compute, storage/data, North-South front-end, and management/OOB. Labels, IL, OTDR, as-builts. Cabling is how those plants are built, not a fifth fabric. This is a factory layer. It is not day-two tidy-up.
East-West compute, and a dedicated storage/data plant where those rails are separate. North-South front-end for users and applications. Management/OOB on its own field, never as application North-South. NVIDIA SuperPOD reference architectures already split compute, storage, and user/NFS onto different switch planes. Structured cabling is how those plants are built.
NP-CBL / Fabric planes
East-West · North-South · Management. Cabling builds the plants.
DWG-CBL-01 · laying out
EW / East-West
Hall-internal scale-out. Compute fabric first. Storage/data fabric on its own plane when dedicated. Not user traffic.
01 / East-West Compute Fabric
MPO/MTP trunks, 400/800G optics, rail-aligned
GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.
02 / Storage / Data Fabric
Separate MPO or LC plant onto the storage leaf when dedicated
Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails.
NS / North-South
Users, applications, API gateway/load balancer, enterprise/core, firewalls. Application-facing Ethernet. Not another GPU rail.
03 / North-South / Front-End Fabric
LC or lower-count MPO, 100G-class typical
Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O.
MGMT / Management
BMC, management, operational control. Never classify as application North-South.
04 / Management / OOB Fabric
Cat6A or dedicated 1/10G optics. Own patch field.
BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark.
DAC, AEC, AOC, MMF, SMF. Pick them for geometry, service, and heat, not for a line-item discount. Campus CIN reach is a fiber number.
NP-CBL-MED / Media and reach
DAC · AEC · AOC · MMF · SMF
DAC
In-rack / adjacent. Typically sub-3 m at 400/800G.
ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend.
Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly.
AEC
Active copper. Longer than DAC, still copper.
In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch.
Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail.
AOC
In-row to in-hall. Fiber with the optics already on the ends.
When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut.
Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant.
MMF (OM4 / OM5)
In-hall SR-class. MPO/MTP trunks.
Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site.
Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested.
SMF (OS2)
Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G.
Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU.
Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone.
NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF.
CIN 500 m class restated from NVIDIA DSX Facilities Infrastructure Reference Design Overview. Public overview only. Not an NVOnline reprint. NeuronPlant is not an NVIDIA Cloud Partner.
Rail-aligned layouts
Each GPU rail lands on the matching leaf. Cables follow the rail map, not the nearest free QSFP. InfiniBand and Ethernet both fail the same way when rails are crossed: collectives look like a software bug.
InfiniBand vs Ethernet cabling
Same cage class does not mean the same plant. Do not mix IB and Ethernet in one trunk. Do not share cassettes. Label the fabric on both ends before the first transceiver seats.
Underfloor vs overhead
High-density GPU halls usually put liquid and power in the floor or the rear, and put fiber in overhead baskets. Underfloor fiber under a liquid loop is how a drip becomes a dark fabric. Pick one primary path. Document the exception.
Cable management
Bend radius is a spec, not a vibe. Slack loops at the rack and at the distribution frame. Strain relief at the NIC, not at the strain of the next installer. Separate baskets per plant so a storage recable cannot yank a compute trunk.
Labeling
Both ends. Unique ID. Fabric, rack, rail, port, polarity type, length. If the label does not match the as-built, the as-built is a drawing, not a plant.
Test
End-face inspection on every MPO. Insertion loss per lane, not per trunk average. Polarity verification before optics. OTDR on SMF backbone. Fail the lane, not the meeting. No test record means the plant is not done.
As-built documentation
Port map, polarity map, tray route, slack locations, test PDFs tied to the label IDs. The next engineer should be able to isolate a rail at 03:00 without calling the person who pulled it.
Bend radius. Dirty MPO. Mixed polarity. No slack. Blocked airflow. Untestable plant. No as-builts. Compute and storage on one tray.
NP-CBL-FAIL / Amateur plant
How a GPU hall stays dark
Bend radius
A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes.
Dirty MPO
One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected.
Mixed polarity
Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks.
No slack
You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage.
Blocked airflow
DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem.
Untestable plant
No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days.
No as-builts
The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name.
Compute and storage on one tray
A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not.
A real unnamed operator bought a large CPU and GPU fleet. Another firm pulled the structured cabling and passive infrastructure. The hall has been offline for a long time. That is not a NeuronPlant delivery. Read the cautionary record.
Cabling is architecture. Trunks, MPO counts, pathways, switch positions, and PDU circuit landings decide whether the next cell is additive or a recable. Pre-build and reserve where the first transition is already visible. Do not pull a Day-1 plant that can only ever be Day-1.
Trunks
Structured trunks with spare fibers, not a jumper for every live port. A spare in the trunk is cheaper than a night in the hall.
MPO counts
Count lanes for the optics class you will grow into, not only the optic seated today. Mixing MPO-12 leftovers into an MPO-16 plant is how the next leaf stays dark.
Pathways
Reserved basket and tray for the next trunk. If the only path is full, additive growth is already a transition.
Switch positions
U and power for the next leaf, with the rail map already drawn. A free ToR slot with no trunk landing is not reserve.
PDUs
Circuit count and landing for the next node or leaf. Cabling the packet path while the power path is exhausted is how a hall dead-ends.
Pre-build / reserve
Justified when Day-1 already names the first transition. Not justified as dark fiber for a fantasy campus. Write the gate. See /expansion.
Pre-build is a Day-1 strategy, not a default. Gates and the money path on Expansion / Day-1. After the plant lands, the path from data to workload still has to be engineered. Applications / Platform.
NP / Intake
One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.