# NeuronPlant / NVIDIA switches Quantum InfiniBand and Spectrum-X Ethernet families, then cited SuperPOD names for those jobs. Canonical HTML: https://neuronplant.com/capabilities#nvidia-switches Collection file: https://neuronplant.com/llms/switch.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # Quantum-X800 Q3400 Fabric: East-West InfiniBand compute Role: Scale-out training / NVL scale-out. East-West compute fabric. NVIDIA SuperPOD B300 RA: 800 Gb/s, rail-optimized fat tree for the compute fabric. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/dgx-superpod-components.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Quantum-2 QM9700 Fabric: InfiniBand compute or storage/data Role: NDR 400 Gb/s leaf. Compute fabric or dedicated storage/data fabric, depending on the SuperPOD job. NVIDIA SuperPOD RA uses MQM9700-NS2R (AC) for InfiniBand storage. Quantum-2 is 64x 400 Gb/s NDR in 1U. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/network-fabrics.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Spectrum-4 SN5600 Fabric: Spectrum-X Ethernet Role: AI Ethernet. Plane depends on the job: East-West compute, storage/data, or North-South front-end. Do not assume one role. 64x 800GbE, Spectrum-4 ASIC. NVIDIA SuperPOD Ethernet RA uses SN5600 / SN5610 for storage and in-band paths. NVIDIA source: https://www.nvidia.com/en-us/networking/spectrumx/ Canonical: https://neuronplant.com/capabilities#nvidia-switches # Spectrum SN5610 Fabric: North-South / user Role: User storage, NFS, homes, application-facing path. Front-end / North-South, not the high-performance storage/data fabric. NVIDIA SuperPOD B300 RA connects user storage to leaf-layer SN5610 (typically 100G-class DR1 for home / metadata). NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/network-fabrics.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Four fabric planes Four planes, not four equal cable plants of the same job. East-West compute fabric is GPU-to-GPU, DGX-to-DGX, scale-out: RDMA/RoCE or InfiniBand, NVIDIA Spectrum-X or Quantum-X. Storage/data is its own internal data plane when physically separate. Do not fold it into a generic East-West line if dedicated. North-South is users, apps, API gateway/LB, enterprise/core. Not another GPU rail. Management/OOB is BMC and operational control. Never application North-South. Converged versus dedicated is a decision on whether compute and storage share a leaf. NVIDIA SuperPOD public RAs already split compute, storage, and user/NFS switch jobs. Structured cabling builds those plants. It is not a fifth fabric. We do not sell a weekly RoCE tune as a service SKU. Congestion behaviour is a plant job: lossless where it is justified, QoS, PFC and ECN as mechanisms, not as a product name. ## Hierarchy EAST-WEST: Compute Fabric; Storage/Data Fabric where dedicated. NORTH-SOUTH: Front-End / Application / Enterprise Fabric. MANAGEMENT: OOB / BMC / Management Fabric. Structured cabling is how those plants are built, not a fifth fabric. ## Topologies - Dedicated planes. When: Collectives and checkpoint burst must not share queues. Training halls, SuperPOD-class jobs, noisy neighbours. Storage/data stays off the East-West compute rails. Cost: More leaves, more trunks, more trays. Cleaner failure domains. Easier to expand one plane without recabling the other. - Converged fabric. When: A smaller cell, a bounded RAG/inference job, or a hall that cannot take a second leaf class. Still name oversub and QoS in writing. Convergence is a compute/storage decision, not permission to mix North-South or OOB onto GPU rails. Cost: Fewer SKUs. Shared congestion and a shared blast radius. The first transition is often splitting the plane you collapsed to save ports. ## Planes to name - East-West / East-West Compute Fabric: Collectives, NCCL, scale-out between nodes or NVL racks. The model runtime lives here. This is not user traffic and not BMC. - East-West / Storage / Data Fabric: Compute to high-performance storage: datasets, checkpoints, RAG index path. Do not fold it into a generic East-West compute line if dedicated. Do not auto-label it North-South or East-West without that context. Decide with the storage machine, not with leftover ports. - North-South / North-South / Front-End Fabric: Application-facing Ethernet. NFS homes, logs, ingest, people. Must not share queues with training I/O. Not just another rail next to the GPU rails. - Management / Management / OOB Fabric: Lights-out, firmware, serial, leak, PDU management if that job exists. In-band plant services sit here, not on application North-South. If this dies inside a compute bundle, the hall is dark and you cannot see why. ## Networking is more than port count - Topology: Leaf/spine, rail alignment, how a new leaf joins. Not a pile of ToR switches. - East-West compute: GPU-to-GPU and DGX-to-DGX diameter, RDMA/RoCE or InfiniBand, Spectrum-X or Quantum-X. The model runtime lives here. Not a generic line that also swallows storage. - Storage / data path: Whether checkpoints and RAG share the compute rails. When dedicated, this is its own internal data plane. Do not auto-label it North-South or East-West without that context. - North-South / front-end: Users, applications, API gateway/LB, enterprise/core. Application-facing Ethernet. Not a GPU rail. - Management / OOB: BMC and operational control. Never classify as application North-South. - Oversubscription: Written per plane. A number that looked fine at eight nodes can fail at the first additive cell. - RoCE / lossless Ethernet: A job when Ethernet carries RDMA on the East-West compute fabric. Design congestion, not a slogan. We do not productise a PFC/ECN recipe as a SKU. - QoS, PFC, ECN: Mechanisms that keep planes from eating each other. Named on the drawing. Tuned in commissioning. Not a catalog line. - Optics: Reach, power, and what the structured cabling can take. 800G on dirty MPO is inventory. - Failure domains: What dies with one leaf, one trunk, one PDU. Blast radius is architecture. - Expansion sequence: Which leaf, which trunk, which reserved position. See /expansion. Port count is not a sequence. Pages: /capabilities#fabrics. Related: /cabling, /expansion, /insights/ethernet-vs-infiniband. Canonical: https://neuronplant.com/capabilities#fabrics