# NeuronPlant / Structured cabling & passive infrastructure Fiber, MPO/MTP, patch panels, pathways, copper, polarity, test records. Canonical HTML: https://neuronplant.com/cabling Collection file: https://neuronplant.com/llms/cabling.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # Structured cabling & passive infrastructure Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. The layer that lands the fabric. Skip it and the GPUs sit dark. ## Four fabric planes (not four equal jobs) EAST-WEST: Compute Fabric, and Storage/Data Fabric where dedicated. NORTH-SOUTH: Front-End / Application / Enterprise Fabric. MANAGEMENT: OOB / BMC / Management Fabric. Structured cabling is how those plants are built, not a fifth fabric. - EAST-WEST / East-West Compute Fabric: MPO/MTP trunks, 400/800G optics, rail-aligned GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray. - EAST-WEST / Storage / Data Fabric: Separate MPO or LC plant onto the storage leaf when dedicated Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails. - NORTH-SOUTH / North-South / Front-End Fabric: LC or lower-count MPO, 100G-class typical Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O. - MANAGEMENT / Management / OOB Fabric: Cat6A or dedicated 1/10G optics. Own patch field. BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark. ## Media - DAC (In-rack / adjacent. Typically sub-3 m at 400/800G.) ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend. Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly. - AEC (Active copper. Longer than DAC, still copper.) In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch. Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail. - AOC (In-row to in-hall. Fiber with the optics already on the ends.) When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut. Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant. - MMF (OM4 / OM5) (In-hall SR-class. MPO/MTP trunks.) Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site. Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested. - SMF (OS2) (Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G.) Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU. Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone. ## Practice - Rail-aligned layouts: Each GPU rail lands on the matching leaf. Cables follow the rail map, not the nearest free QSFP. InfiniBand and Ethernet both fail the same way when rails are crossed: collectives look like a software bug. - InfiniBand vs Ethernet cabling: Same cage class does not mean the same plant. Do not mix IB and Ethernet in one trunk. Do not share cassettes. Label the fabric on both ends before the first transceiver seats. - Underfloor vs overhead: High-density GPU halls usually put liquid and power in the floor or the rear, and put fiber in overhead baskets. Underfloor fiber under a liquid loop is how a drip becomes a dark fabric. Pick one primary path. Document the exception. - Cable management: Bend radius is a spec, not a vibe. Slack loops at the rack and at the distribution frame. Strain relief at the NIC, not at the strain of the next installer. Separate baskets per plant so a storage recable cannot yank a compute trunk. - Labeling: Both ends. Unique ID. Fabric, rack, rail, port, polarity type, length. If the label does not match the as-built, the as-built is a drawing, not a plant. - Test: End-face inspection on every MPO. Insertion loss per lane, not per trunk average. Polarity verification before optics. OTDR on SMF backbone. Fail the lane, not the meeting. No test record means the plant is not done. - As-built documentation: Port map, polarity map, tray route, slack locations, test PDFs tied to the label IDs. The next engineer should be able to isolate a rail at 03:00 without calling the person who pulled it. ## Why amateur cabling takes a GPU hall off the air - Bend radius: A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes. - Dirty MPO: One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected. - Mixed polarity: Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks. - No slack: You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage. - Blocked airflow: DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem. - Untestable plant: No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days. - No as-builts: The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name. - Compute and storage on one tray: A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not. ## CIN reach NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF. ## Ports on the switch are not a growth plan. Cabling is architecture. Trunks, MPO counts, pathways, switch positions, and PDU circuit landings decide whether the next cell is additive or a recable. Pre-build and reserve where the first transition is already visible. Do not pull a Day-1 plant that can only ever be Day-1. - Trunks: Structured trunks with spare fibers, not a jumper for every live port. A spare in the trunk is cheaper than a night in the hall. - MPO counts: Count lanes for the optics class you will grow into, not only the optic seated today. Mixing MPO-12 leftovers into an MPO-16 plant is how the next leaf stays dark. - Pathways: Reserved basket and tray for the next trunk. If the only path is full, additive growth is already a transition. - Switch positions: U and power for the next leaf, with the rail map already drawn. A free ToR slot with no trunk landing is not reserve. - PDUs: Circuit count and landing for the next node or leaf. Cabling the packet path while the power path is exhausted is how a hall dead-ends. - Pre-build / reserve: Justified when Day-1 already names the first transition. Not justified as dark fiber for a fantasy campus. Write the gate. See /expansion. Canonical: https://neuronplant.com/cabling # East-West Compute Fabric Band: east-west MPO/MTP trunks, 400/800G optics, rail-aligned GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray. Canonical: https://neuronplant.com/cabling#plants # Storage / Data Fabric Band: east-west Separate MPO or LC plant onto the storage leaf when dedicated Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails. Canonical: https://neuronplant.com/cabling#plants # North-South / Front-End Fabric Band: north-south LC or lower-count MPO, 100G-class typical Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O. Canonical: https://neuronplant.com/cabling#plants # Management / OOB Fabric Band: management Cat6A or dedicated 1/10G optics. Own patch field. BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark. Canonical: https://neuronplant.com/cabling#plants # DAC In-rack / adjacent. Typically sub-3 m at 400/800G. ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend. Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly. Canonical: https://neuronplant.com/cabling#media # AEC Active copper. Longer than DAC, still copper. In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch. Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail. Canonical: https://neuronplant.com/cabling#media # AOC In-row to in-hall. Fiber with the optics already on the ends. When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut. Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant. Canonical: https://neuronplant.com/cabling#media # MMF (OM4 / OM5) In-hall SR-class. MPO/MTP trunks. Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site. Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested. Canonical: https://neuronplant.com/cabling#media # SMF (OS2) Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G. Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU. Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone. Canonical: https://neuronplant.com/cabling#media # Rail-aligned layouts Each GPU rail lands on the matching leaf. Cables follow the rail map, not the nearest free QSFP. InfiniBand and Ethernet both fail the same way when rails are crossed: collectives look like a software bug. Canonical: https://neuronplant.com/cabling#practice # InfiniBand vs Ethernet cabling Same cage class does not mean the same plant. Do not mix IB and Ethernet in one trunk. Do not share cassettes. Label the fabric on both ends before the first transceiver seats. Canonical: https://neuronplant.com/cabling#practice # Underfloor vs overhead High-density GPU halls usually put liquid and power in the floor or the rear, and put fiber in overhead baskets. Underfloor fiber under a liquid loop is how a drip becomes a dark fabric. Pick one primary path. Document the exception. Canonical: https://neuronplant.com/cabling#practice # Cable management Bend radius is a spec, not a vibe. Slack loops at the rack and at the distribution frame. Strain relief at the NIC, not at the strain of the next installer. Separate baskets per plant so a storage recable cannot yank a compute trunk. Canonical: https://neuronplant.com/cabling#practice # Labeling Both ends. Unique ID. Fabric, rack, rail, port, polarity type, length. If the label does not match the as-built, the as-built is a drawing, not a plant. Canonical: https://neuronplant.com/cabling#practice # Test End-face inspection on every MPO. Insertion loss per lane, not per trunk average. Polarity verification before optics. OTDR on SMF backbone. Fail the lane, not the meeting. No test record means the plant is not done. Canonical: https://neuronplant.com/cabling#practice # As-built documentation Port map, polarity map, tray route, slack locations, test PDFs tied to the label IDs. The next engineer should be able to isolate a rail at 03:00 without calling the person who pulled it. Canonical: https://neuronplant.com/cabling#practice # Bend radius A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes. Canonical: https://neuronplant.com/cabling#failure # Dirty MPO One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected. Canonical: https://neuronplant.com/cabling#failure # Mixed polarity Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks. Canonical: https://neuronplant.com/cabling#failure # No slack You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage. Canonical: https://neuronplant.com/cabling#failure # Blocked airflow DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem. Canonical: https://neuronplant.com/cabling#failure # Untestable plant No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days. Canonical: https://neuronplant.com/cabling#failure # No as-builts The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name. Canonical: https://neuronplant.com/cabling#failure # Compute and storage on one tray A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not. Canonical: https://neuronplant.com/cabling#failure # CIN reach NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF. Canonical: https://neuronplant.com/cabling#media # Ports on the switch are not a growth plan. Cabling is architecture. Trunks, MPO counts, pathways, switch positions, and PDU circuit landings decide whether the next cell is additive or a recable. Pre-build and reserve where the first transition is already visible. Do not pull a Day-1 plant that can only ever be Day-1. - Trunks: Structured trunks with spare fibers, not a jumper for every live port. A spare in the trunk is cheaper than a night in the hall. - MPO counts: Count lanes for the optics class you will grow into, not only the optic seated today. Mixing MPO-12 leftovers into an MPO-16 plant is how the next leaf stays dark. - Pathways: Reserved basket and tray for the next trunk. If the only path is full, additive growth is already a transition. - Switch positions: U and power for the next leaf, with the rail map already drawn. A free ToR slot with no trunk landing is not reserve. - PDUs: Circuit count and landing for the next node or leaf. Cabling the packet path while the power path is exhausted is how a hall dead-ends. - Pre-build / reserve: Justified when Day-1 already names the first transition. Not justified as dark fiber for a fantasy campus. Write the gate. See /expansion. Canonical: https://neuronplant.com/cabling#tomorrow