U / rack space
Necessary and almost never sufficient. Count the accelerators, CDU manifolds, patch, and service U. Then ask whether power and heat still fit.
Capabilities are the evidence behind a NeuronPlant delivery, organized by accelerator-plant layer. Site, power, and cooling are constraints of the hall we occupy. Each layer is delivered as part of one project, not as a standalone product line, and not as a data-center general contract.
01
We read halls the way a high-density cell is read: floor load, aisle strategy, containment, intake path, and whether the room can accept the next density jump without a rebuild. U left is not power left. Those numbers are the hall's. Facility is architecture of the room we place into, not NeuronPlant as the data-center general contractor.
A hall we occupy is a set of coupled limits: U, weight, floor load, normal and peak kW, A/B, PDU, heat, clearance, and pathways. Empty U with a dead feed is not capacity. The hall owner supplies the room. NeuronPlant engineers the GPU / TPU / ASIC cell to fit, or matches another hall. We do not pour the slab.
U / rack space
Necessary and almost never sufficient. Count the accelerators, CDU manifolds, patch, and service U. Then ask whether power and heat still fit.
Weight
Liquid-cooled GPU racks and full storage frames are plant loads. If the slab or raised floor cannot take them, the SKU does not ship, or it ships and sits.
Floor load
kN/m², not a vibe. Point loads under NVL-class racks and CDUs. Path in from the dock matters as much as the pad.
Normal and peak kW
Nameplate is not the bill. Training peaks. Inference is different. Design the feed, UPS, and breaker to the number that shows up, A/B where the architecture requires it.
A/B
Dual cord is a drawing: two PDUs, two paths, maintenance bypass. Dual cord on a single upstream is a sticker.
PDU
Circuit count, phase, and whether managed is justified. See Power / PDU. Reserve circuits for the additive cell or admit the row is full.
Heat
Rejection path, not a CRAH brochure. Air until physics fails. Then water, CDU, and isolation. Peak kW is peak heat.
Clearance
Service, manifolds, and a sled that can come out without tensioning fiber. A packed hot aisle is not density. It is an unmaintainable plant.
Pathways
Fiber baskets, power, liquid. Overhead versus underfloor. The next trunk needs a reserved path. See Cabling / tomorrow.
02
From the feed the hall presents to the rack PDU. Capacity, redundancy topology, maintenance bypass, and the difference between a nameplate and a sustained training load. Managed vs unmanaged is a constraint: specify distribution, metering, or outlet switching. Do not default to a management plane. Utility interconnection is a hall-owner or custom-system-supplier job.
A PDU is how the rack gets power. Managed means the PDU has a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Managed is not automatic. Specify the job: unmanaged distribution, metered, or switched at the outlet.
This is not a vendor catalog. Managed PDU and unmanaged PDU are the durable words.
NP-PWR-PDU / Dual-cord rack
Unmanaged · metered · switched. Specify the job.
01 / Unmanaged distribution
Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks.
02 / Metered
See kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. Justified when the board or BMS does not already meter at this grain, or when the hall is multi-tenant or billed.
03 / Switched / outlet-level
Remote outlet control. On dual-cord GPU and compute racks this is a foot-gun: you do not power-cycle a training node from a PDU web UI. Specify it only when that job is real.
When managed is justified
When you do not need managed
03
Air until it fails the physics of this room. Then facility water, CDU, manifold, and rejection, engineered to what the hall can present, with isolation and leak response designed before the first cold plate is specified.
04
The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. We do not maintain a weekly chip catalog. NVIDIA DGX B200 and NVIDIA DGX B300 for 8-GPU nodes. NVIDIA DGX GB200 NVL72 and NVIDIA DGX GB300 NVL72 when the model needs a 72-GPU NVLink domain. HGX OEM nodes where the hall already standardizes on that tray. Specs are NVIDIA's; the room we occupy, the loop and fabric are the design.
05
NVIDIA Quantum InfiniBand and Spectrum-X Ethernet. Four planes, not four equal jobs: East-West compute fabric (GPU-to-GPU, DGX-to-DGX, RDMA/RoCE or InfiniBand), storage/data fabric when dedicated, North-South / front-end for users and apps, management/OOB never as application North-South. Compute versus storage is a decision: converged or dedicated. SuperPOD reference architectures name jobs on those planes. Cited names sit on the cards below, not as a weekly catalog. Networking is topology, oversub, optics, and failure domains, not port count.
06
Structured fiber and copper as a factory layer. MMF/SMF, MTP/MPO, 400/800G optics, DAC/AEC/AOC, rail-aligned InfiniBand and Ethernet. Trays grouped by plane: East-West compute, storage/data, North-South front-end, management/OOB. Cabling is how those plants are built, not a fifth fabric. Labels, insertion loss, OTDR, polarity, as-builts. Amateur plant is how a GPU hall stays dark.
07
Capacity is not a spec. Throughput, IOPS, metadata, checkpoints, RAG, and protection first, then the product. Training looks for sequential GB/s and checkpoint burst (NVIDIA SuperPOD: >40 GB/s per NVIDIA DGX B300 node, >80 GB/s on the XDR RA). RAG is sized from source data, not vector brochure TB. Inference looks for local NIM weights, and on a large serving job a KV / prefix cache hierarchy (HBM → RAM → NVMe → remote). Those are different machines.
08
The plant is not finished when the accelerators power on. We size the path from sources through ingest, storage, train or RAG, accelerator runtime, and inference serving. On a large serving job, KV and prefix cache are a memory hierarchy, not another storage array. NVIDIA NIM profiles (precision, tensor-parallel, LoRA) come from NVIDIA's public support matrix. We integrate that serving plane onto the factory. We do not claim to write the customer's models.
09
After the platform exists, the factory still has to schedule and stay visible. Docker is the image. Kubernetes is a starting cell of three masters and three workers. Masters may be VMs. Workers that hold accelerators are usually bare metal. Observe is part of that plane. Fluent Bit is the example log shipper, not a dashboard catalog. Full cell on Runtime / Operate.
East-West compute fabric is GPU-to-GPU, DGX-to-DGX, scale-out: RDMA/RoCE or InfiniBand, NVIDIA Spectrum-X or Quantum-X. Storage/data is its own internal data plane when physically separate. Do not fold it into a generic East-West line if dedicated. North-South is users, apps, API gateway/LB, enterprise/core. Not another GPU rail. Management/OOB is BMC and operational control. Never application North-South. Converged versus dedicated is a decision on whether compute and storage share a leaf. NVIDIA SuperPOD public RAs already split compute, storage, and user/NFS switch jobs. Structured cabling builds those plants. It is not a fifth fabric.
We do not sell a weekly RoCE tune as a service SKU. Congestion behaviour is a plant job: lossless where it is justified, QoS, PFC and ECN as mechanisms, not as a product name.
NP-NET / Fabric decision
Not one true topology. East-West compute, storage/data where dedicated, North-South front-end, and management/OOB still have to be written down. They are not four sibling jobs of equal role.
01 / Dedicated planes
Collectives and checkpoint burst must not share queues. Training halls, SuperPOD-class jobs, noisy neighbours. Storage/data stays off the East-West compute rails.
More leaves, more trunks, more trays. Cleaner failure domains. Easier to expand one plane without recabling the other.
02 / Converged fabric
A smaller cell, a bounded RAG/inference job, or a hall that cannot take a second leaf class. Still name oversub and QoS in writing. Convergence is a compute/storage decision, not permission to mix North-South or OOB onto GPU rails.
Fewer SKUs. Shared congestion and a shared blast radius. The first transition is often splitting the plane you collapsed to save ports.
East-West
01 / East-West Compute Fabric
Collectives, NCCL, scale-out between nodes or NVL racks. The model runtime lives here. This is not user traffic and not BMC.
02 / Storage / Data Fabric
Compute to high-performance storage: datasets, checkpoints, RAG index path. Do not fold it into a generic East-West compute line if dedicated. Do not auto-label it North-South or East-West without that context. Decide with the storage machine, not with leftover ports.
North-South
03 / North-South / Front-End Fabric
Application-facing Ethernet. NFS homes, logs, ingest, people. Must not share queues with training I/O. Not just another rail next to the GPU rails.
Management
04 / Management / OOB Fabric
Lights-out, firmware, serial, leak, PDU management if that job exists. In-band plant services sit here, not on application North-South. If this dies inside a compute bundle, the hall is dark and you cannot see why.
Topology
Leaf/spine, rail alignment, how a new leaf joins. Not a pile of ToR switches.
East-West compute
GPU-to-GPU and DGX-to-DGX diameter, RDMA/RoCE or InfiniBand, Spectrum-X or Quantum-X. The model runtime lives here. Not a generic line that also swallows storage.
Storage / data path
Whether checkpoints and RAG share the compute rails. When dedicated, this is its own internal data plane. Do not auto-label it North-South or East-West without that context.
North-South / front-end
Users, applications, API gateway/LB, enterprise/core. Application-facing Ethernet. Not a GPU rail.
Management / OOB
BMC and operational control. Never classify as application North-South.
Oversubscription
Written per plane. A number that looked fine at eight nodes can fail at the first additive cell.
RoCE / lossless Ethernet
A job when Ethernet carries RDMA on the East-West compute fabric. Design congestion, not a slogan. We do not productise a PFC/ECN recipe as a SKU.
QoS, PFC, ECN
Mechanisms that keep planes from eating each other. Named on the drawing. Tuned in commissioning. Not a catalog line.
Optics
Reach, power, and what the structured cabling can take. 800G on dirty MPO is inventory.
Failure domains
What dies with one leaf, one trunk, one PDU. Blast radius is architecture.
Expansion sequence
Which leaf, which trunk, which reserved position. See /expansion. Port count is not a sequence.
Cited SuperPOD switch names below. Expansion sequence on Expansion. Structured cabling on Structured cabling.
Families, then cited SuperPOD names. Not a weekly SKU catalog.
NP-NET / NVIDIA switch planes
East-West · North-South · Management. Families, not SKUs.
East-West
01
NVIDIA Quantum-X / Spectrum-X
RDMA/RoCE or InfiniBand. GPU-to-GPU, DGX-to-DGX, scale-out.
800 Gb/s XDR / 400 Gb/s NDR / 800GbE class
Collectives, NCCL, scale-out between nodes or NVL racks. The model runtime lives here. This is not user traffic and not BMC.
02
Quantum or Spectrum-X
Own data plane when dedicated. Off the compute rails.
400 Gb/s IB or 800GbE
Compute to high-performance storage: datasets, checkpoints, RAG index path. Do not fold it into a generic East-West compute line if dedicated. Do not auto-label it North-South or East-West without that context. Decide with the storage machine, not with leftover ports.
North-South
03
Spectrum-X Ethernet
Users, apps, API gateway/LB, enterprise/core, firewalls.
100G-class typical
Application-facing Ethernet. NFS homes, logs, ingest, people. Must not share queues with training I/O. Not just another rail next to the GPU rails.
Management
04
Dedicated 1/10G / Cat6A
BMC, management, operational control. Own patch field.
1/10G class
Lights-out, firmware, serial, leak, PDU management if that job exists. In-band plant services sit here, not on application North-South. If this dies inside a compute bundle, the hall is dark and you cannot see why.
DWG-NET-01 · laying out
East-West InfiniBand compute
Scale-out training / NVL scale-out. East-West compute fabric.
NVIDIA SuperPOD B300 RA: 800 Gb/s, rail-optimized fat tree for the compute fabric.
InfiniBand compute or storage/data
NDR 400 Gb/s leaf. Compute fabric or dedicated storage/data fabric, depending on the SuperPOD job.
NVIDIA SuperPOD RA uses MQM9700-NS2R (AC) for InfiniBand storage. Quantum-2 is 64x 400 Gb/s NDR in 1U.
Spectrum-X Ethernet
AI Ethernet. Plane depends on the job: East-West compute, storage/data, or North-South front-end. Do not assume one role.
64x 800GbE, Spectrum-4 ASIC. NVIDIA SuperPOD Ethernet RA uses SN5600 / SN5610 for storage and in-band paths.
North-South / user
User storage, NFS, homes, application-facing path. Front-end / North-South, not the high-performance storage/data fabric.
NVIDIA SuperPOD B300 RA connects user storage to leaf-layer SN5610 (typically 100G-class DR1 for home / metadata).
Equal weight to the switch planes. Fiber, MPO/MTP trunks, patch panels, pathways, copper, and physical connectivity reserved for expansion. Structured cabling is how East-West, North-South, and management plants are built, not a fifth fabric. The GPUs do not come up if this is amateur.
NP-CBL / Fabric planes
East-West · North-South · Management. Cabling builds the plants.
DWG-CBL-01 · laying out
EW / East-West
Hall-internal scale-out. Compute fabric first. Storage/data fabric on its own plane when dedicated. Not user traffic.
01 / East-West Compute Fabric
MPO/MTP trunks, 400/800G optics, rail-aligned
GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.
02 / Storage / Data Fabric
Separate MPO or LC plant onto the storage leaf when dedicated
Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails.
NS / North-South
Users, applications, API gateway/load balancer, enterprise/core, firewalls. Application-facing Ethernet. Not another GPU rail.
03 / North-South / Front-End Fabric
LC or lower-count MPO, 100G-class typical
Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O.
MGMT / Management
BMC, management, operational control. Never classify as application North-South.
04 / Management / OOB Fabric
Cat6A or dedicated 1/10G optics. Own patch field.
BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark.
NP-CBL-FAIL / Amateur plant
How a GPU hall stays dark
Bend radius
A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes.
Dirty MPO
One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected.
Mixed polarity
Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks.
No slack
You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage.
Blocked airflow
DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem.
Untestable plant
No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days.
No as-builts
The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name.
Compute and storage on one tray
A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not.
Full specification on Structured Cabling & Passive Infrastructure, including the plant built for tomorrow.
NP-STO / What to look for
Throughput numbers below for training come from the NVIDIA SuperPOD with NVIDIA DGX B300 reference architectures. RAG and NIM notes follow how those stacks actually move bytes, plus the NVIDIA NIM support matrix for profile sizing.
Training / post-training
Sustained sequential GB/s and checkpoint burst, not IOPS marketing.
Shared parallel filesystem for datasets. Separate checkpoint target that can absorb a simultaneous flush from the job. NVIDIA SuperPOD B300 RA states high-performance storage I/O per node must exceed 40 GB/s (Ethernet RA) and, on the Quantum-X800 RA, 80 GB/s.
Typical: HPS on IB (QM9700) or RoCE (SN5600) as the storage/data fabric, plus a quieter NFS/user tier on SN5610 as North-South / front-end. Do not fold HPS into a generic East-West compute line if dedicated.
Avoid: A single enterprise NAS asked to be both home directories and the training scratch.
RAG / enterprise retrieval
Random read IOPS, metadata, and a clean split between documents and the vector index.
Object or file store for source documents. NVMe-backed or in-memory vector index. Ingest path sized to nightly rebuilds, not to training checkpoints.
Typical: Object + vector engine. North-South / front-end Ethernet is usually enough if the index is local or on a low-latency tier. That is not the East-West compute fabric.
Avoid: Putting the vector store on the same parallel FS as a training job.
Inference / NIM serving
Fast model-weight load, local NVMe or a warm cache, isolation from training I/O, and a KV / prefix cache plan on large serving.
NVIDIA NIM selects a profile from its support matrix (precision, tensor-parallel size, LoRA). Weights load from local NVMe. That is not the whole inference memory story. On a large workload, KV cache needs a hierarchy: GPU HBM, host overflow, and sometimes a shared prefix cache across the serving plane. That tier is not a storage array and not a checkpoint pool.
Typical: Node-local NVMe for NIM images and weights. KV in HBM first. Host memory and a serving-plane prefix cache when context and concurrency fill the GPU.
Avoid: Loading 70B-class weights over a contended NFS home share. Treating KV as another LUN on the training filesystem.
Research / HPC mix
Scratch that can be purged, project space that cannot, and a scheduler that knows the difference.
Burst scratch on the HPS fabric. Project datasets on a capacity tier. Home and logs on the user-storage Ethernet path, as in the NVIDIA SuperPOD split.
Typical: Two storage systems: HPS + user storage. NVIDIA documents this split explicitly.
Avoid: One quota, one protocol, every lab on the same queue.
TB on a quote is not a spec. Engineer the requirement before the product: throughput, IOPS, metadata, checkpoints, RAG, and protection. Then pick a machine that can do that job. Training, RAG, and inference remain different machines.
Throughput
Sequential GB/s for training scratch and checkpoint burst. NVIDIA SuperPOD public RAs already publish per-node floors. Size to the job, not to a single array's headline.
IOPS
Random read for RAG retrieve and metadata-heavy ingest. A parallel FS built for sequential train I/O is the wrong machine.
Metadata
Small files, many objects, index catalogs. The namespace can saturate before the bytes do.
Checkpoints
A simultaneous flush is a storage event. Burst target and recovery window are part of the plant, not an afterthought quota.
RAG
Source corpus, embeddings, and the index are different sizes and different I/O. See RAG storage below. Do not size from vector capacity alone.
Protection
Copies, erasure, snapshots, rebuild time. Protection is throughput you no longer have for the job while a disk is gone.
Start with raw TB and how it grows. Then decide copy versus reference, retention, embedding volume, protection, and what a re-index costs. The vector store is one downstream number. It is not the architecture.
Raw TB
Authoritative source size today. Everything else is derived.
Growth
Ingest rate and how long the plant must take it without a storage transition.
Copy vs reference
Landing zones that copy the corpus multiply TB. Reference in place saves bytes and couples failure domains. Write the choice.
Retention
How long each data state lives. Legal hold is not the same as hot retrieve.
Embeddings
A function of chunking, dimensions, and versions kept. A new embedding model is a storage event, not a config flag.
Protection
Which states are rebuilt from source, which are backed up, which are disposable.
Re-index
Bytes, IOPS, and window to rebuild the index from chunks. If you cannot re-index, you do not own the RAG plant.
Large serving also has a KV / prefix cache tier. That is a serving-plane memory hierarchy, not a storage array. Definition on Applications / Platform. RAG path and rebuildability on Applications.
Serving, RAG services, NIM / runtime integration, and the ingest path belong on the plant. Not as a generic software product.
NP-PIPE / Data to workload
Schematic · not a product flowchart
Full path on Applications / Platform. Docker, Kubernetes, the 3+3 cell, and Fluent Bit on Runtime / Operate. KV / prefix cache on Applications.
NP / Intake
One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.