# NeuronPlant > We Build AI Factories. One Partner. Design Authority and Delivery. The Entire AI Factory. This file is a short map. Full technical corpora live at /llms/{collection}.txt, on the HTML topic pages, and in /llms-full.txt organized by collection. Prefer MCP or JSON over scraping HTML. NeuronPlant is the design authority and the single accountable delivery partner for an AI factory. You bring the AI requirement. NeuronPlant engineers the architecture, can purchase the required equipment against that architecture, coordinates suppliers, integrates, validates, and hands over the factory. We do not build the data center. The hall operator supplies and operates the room. NeuronPlant is the design authority and the single accountable delivery partner for an AI factory. NeuronPlant engineers the architecture, can purchase the required equipment against that architecture, coordinates suppliers, and delivers the GPU, TPU, and ASIC plant. The hall operator supplies and operates the room. Sequence: Site → Power → Cooling → Compute → Network → Structured cabling & passive infrastructure → Storage → Platform → Operate. Site in that line is the hall we occupy, not a campus we pour. Not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval (same gate as /customers). ## MCP - [MCP server](https://neuronplant.com/api/mcp): Streamable HTTP, read-only, no auth. Cursor config is on https://neuronplant.com/mcp. - [MCP connect page](https://neuronplant.com/mcp): tools, resources, and client snippets. - [MCP server card](https://neuronplant.com/.well-known/mcp.json) - [Structured JSON](https://neuronplant.com/api/knowledge): full catalog index. ## Site map - [Homepage](https://neuronplant.com/): Buying-story homepage. From AI requirement to a delivered AI factory. One customer journey: Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover. Nine plant layers are canonical. System view lives on /what-we-build#system-view, unnumbered. Header actions: Start a Project and Become a Partner. - [What we build](https://neuronplant.com/what-we-build): Nine plant layers of the accelerator cell including structured cabling & passive infrastructure, the AI platform, operate / monitor, and Expansion / Day-1. Site is the hall we occupy. - [Capabilities](https://neuronplant.com/capabilities): How each plant layer is delivered. Facility as constraints of the hall we occupy, rack PDUs as a spec job, fabric decision (converged vs dedicated), Quantum and Spectrum-X families with cited SuperPOD names, storage sized as a requirement, operate / monitor, and cited model profiles for sizing. Not NP as a data-center general contractor. - [AI factories](https://neuronplant.com/ai-factories): Factory templates NeuronPlant can deliver for private AI, training, AI-as-a-Service, sovereign, and hall-scale plants, plus how we read the public NVIDIA DSX campus map. - [Models](https://neuronplant.com/models): NVIDIA-cited model and platform references from the public NIM support matrix. - [NVIDIA DSX](https://neuronplant.com/dsx): How we read public NVIDIA DSX campus docs when a hall partner or customer program needs that language. NVOnline packs named only. No partnership claim. Not a campus NeuronPlant builds. - [Structured cabling & passive infrastructure](https://neuronplant.com/cabling): Structured fiber, copper, polarity, test records, and the plant built for tomorrow. - [Growth is an architecture evolution, not a node count.](https://neuronplant.com/expansion): Adding capacity inside a designed envelope is additive. Changing the fabric, the cooling loop, the storage machine, or the hall is a transition. Day-1 names those points. It does not promise unlimited scale. - [Applications / Platform](https://neuronplant.com/applications): AI platform layer after the accelerators exist: serving, RAG services, NIM / runtime integration, and the data path from sources to workload. Large serving includes a KV hierarchy (HBM → RAM → NVMe → remote). Runtime / operate is Docker, Kubernetes, and observe. - [Runtime / Operate](https://neuronplant.com/runtime): Docker images, Kubernetes 3+3 cell, VM-capable masters, production plane (CPU vs accelerator workers, registry, secrets, identity, RAG services), Fluent Bit as the example log shipper. Durable names, not a SKU catalog. - [Projects](https://neuronplant.com/projects): Factory records by sector. Includes a cautionary structured-cabling failure, not a win. - [Customers](https://neuronplant.com/customers): Named only after written approval. Until then, sector and region. - [Insights](https://neuronplant.com/insights): Short technical notes. Constraints first. - [About](https://neuronplant.com/about): Why the firm exists: design authority and single accountable delivery partner from AI requirement to a delivered factory. Not a DC contractor. Plus the integrity stance. Company collection also holds partners, partner intake, and customers. - [Partners](https://neuronplant.com/partners): Two kinds of partners behind a NeuronPlant delivery. Technology: OEMs and platforms that make the accelerator plant (the stack). Delivery: integrators, operators, and hall/colo suppliers who provide the room (who ships a layer). The customer still buys one project from NeuronPlant. NeuronPlant is not the colo. No NVIDIA or Lenovo logo wall. No NCP claim. - [Become a partner](https://neuronplant.com/become-a-partner): Engineering intake to qualify as Technology or Delivery, then by layer (including hall), supply, geography, and NVIDIA/DSX relationship. - [Client portal preview](https://neuronplant.com/portal): Preview of the customer's project environment: status, architecture, facility, suppliers, procurement, milestones, documentation, and handover. Not live. Does not authenticate. Not the product. Not a hall marketplace. No invented customers or project records. - [MCP server](https://neuronplant.com/mcp): How to plug the published factory record into Cursor, Claude Desktop, and other MCP clients. - [llms.txt](https://neuronplant.com/llms.txt): Machine-readable index of routes and knowledge collections for agents. ## Knowledge collections - [Site routes](https://neuronplant.com/llms/route.txt): Important pages and what they contain. (https://neuronplant.com/) - [Plant layers](https://neuronplant.com/llms/system.txt): Site, power, cooling, compute (GPU, TPU, other ASICs), network, structured cabling & passive infrastructure, storage, platform, operate. (https://neuronplant.com/what-we-build) - [NVIDIA platforms](https://neuronplant.com/llms/platform.txt): NVIDIA DGX B200 / NVIDIA DGX B300 nodes and NVIDIA DGX GB200 NVL72 / NVIDIA DGX GB300 NVL72 racks, cited from public NVIDIA pages. (https://neuronplant.com/models) - [NVIDIA switches](https://neuronplant.com/llms/switch.txt): Quantum InfiniBand and Spectrum-X Ethernet families, then cited SuperPOD names for those jobs. (https://neuronplant.com/capabilities#nvidia-switches) - [Storage by workload](https://neuronplant.com/llms/storage.txt): Training, RAG, inference, and HPC mixes as different machines. Large inference also has a KV / prefix cache tier. (https://neuronplant.com/capabilities#storage) - [DSX campus](https://neuronplant.com/llms/dsx.txt): How we read public NVIDIA DSX campus layers for a hall partner or customer program. Not a campus NeuronPlant builds. NVOnline packs named only, not reprinted. (https://neuronplant.com/dsx) - [Factory templates](https://neuronplant.com/llms/factory.txt): Private AI, training, AI-as-a-Service, sovereign, and large-scale factories NeuronPlant can deliver. (https://neuronplant.com/ai-factories) - [Design decisions](https://neuronplant.com/llms/decision.txt): Why this fabric, cooling, storage, sequence, or plant. Rationale in writing. (https://neuronplant.com/mcp) - [Plant notes](https://neuronplant.com/llms/insight.txt): Short technical notes on NVL72, power/cooling, fabrics, structured cabling, and steered RFPs. (https://neuronplant.com/insights) - [Cited models](https://neuronplant.com/llms/model.txt): NVIDIA NIM support-matrix profiles used to size nodes vs racks. (https://neuronplant.com/models) - [Structured cabling & passive infrastructure](https://neuronplant.com/llms/cabling.txt): Fiber, MPO/MTP, patch panels, pathways, copper, polarity, test records. (https://neuronplant.com/cabling) - [AI platform](https://neuronplant.com/llms/pipeline.txt): Serving, RAG services, NIM / runtime integration, and the data path from sources through ingest, train or RAG, Docker/Kubernetes, KV / prefix cache, operate / observe, handover. (https://neuronplant.com/applications) - [Delivery records](https://neuronplant.com/llms/project.txt): Anonymized factory records. No invented customer names. (https://neuronplant.com/projects) - [Customer specifications](https://neuronplant.com/llms/customer-spec.txt): Held behind the same written-approval gate as named customers. (https://neuronplant.com/customers) ## Optional - [llms-full.txt](https://neuronplant.com/llms-full.txt): concatenated published knowledge, grouped by collection. - [NVIDIA public sources](https://neuronplant.com/models): cited, not restated as NeuronPlant measurements. - Held customer specs: sector and region only until written approval. # Published records by collection ## Plant layers Site, power, cooling, compute (GPU, TPU, other ASICs), network, structured cabling & passive infrastructure, storage, platform, operate. HTML: https://neuronplant.com/what-we-build Split file: https://neuronplant.com/llms/system.txt # Site The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it. - Hall / colo cage we place into - Partner room or custom-system supplier - Rack planning in that room - Floor loading - U / weight / clearance - kW and heat the hall can present - Path in / meet-me Canonical: https://neuronplant.com/what-we-build#site # Power From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job. - Hall-presented power - Capacity planning for the cell - UPS as the room allows - PDU / busway in the cell - Managed vs unmanaged PDU - Rack power distribution - Redundancy architecture Canonical: https://neuronplant.com/what-we-build#power # Cooling Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies. - Liquid cooling - CDU design - Facility water requirements - Heat rejection - Chillers - Cooling redundancy Canonical: https://neuronplant.com/what-we-build#cooling # Compute Accelerators as an engineered cluster. GPU is the common case, not the only one. The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. NVIDIA DGX and GB-class remain how we size a GPU plant. Specs are NVIDIA's. The room we occupy, the loop, and the fabric are the design. We do not maintain a weekly chip catalog. - GPU (common case) - TPU - Other accelerators / ASICs - NVIDIA DGX / GB-class when the plant is GPU - Cluster architecture - Capacity planning Canonical: https://neuronplant.com/what-we-build#compute # Network East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South. - NVIDIA Quantum-X / Spectrum-X East-West compute - Storage / data fabric (dedicated when required) - North-South / front-end fabric - Management / OOB fabric - Converged vs dedicated compute and storage - 800G fabrics Canonical: https://neuronplant.com/what-we-build#network # Structured cabling & passive infrastructure Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark. - Fiber and MPO/MTP trunks (MMF / SMF, 400/800G optics) - Patch panels, cross-connects, rack-to-rack - Cable pathways and cable management - Copper and passive optical infrastructure - Separate trays for EW compute, storage/data, N-S, and OOB - Physical connectivity reserved for expansion - Labeling, IL / OTDR, polarity test, as-builts Canonical: https://neuronplant.com/cabling # Storage Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure. - High-performance storage - Throughput / IOPS / metadata - AI datasets - Training storage - RAG sized from source data - Checkpointing - Backup and data lifecycle Canonical: https://neuronplant.com/what-we-build#storage # Platform The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model. - Platform architecture - Ingest and dataset path - Train / fine-tune vs RAG - Accelerator runtime placement - NVIDIA NIM profile sizing - KV / prefix cache (large serving) - Orchestration / MLOps interfaces - Enterprise application handover Canonical: https://neuronplant.com/applications # Operate The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog. - Docker images - Kubernetes cell (3 masters + 3 workers) - Masters may be VMs - CPU workers vs accelerator workers - Observe / monitor - Fluent Bit (example shipper) - Runbooks after handover Canonical: https://neuronplant.com/runtime # Rack PDU: managed vs unmanaged Managed PDU is a spec job, not a default. A PDU is how the rack gets power. Managed means the PDU has a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Managed is not automatic. Specify the job: unmanaged distribution, metered, or switched at the outlet. This is not a vendor catalog. Managed PDU and unmanaged PDU are the durable words. ## Jobs - 01 Unmanaged distribution: Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks. - 02 Metered: See kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. Justified when the board or BMS does not already meter at this grain, or when the hall is multi-tenant or billed. - 03 Switched / outlet-level: Remote outlet control. On dual-cord GPU and compute racks this is a foot-gun: you do not power-cycle a training node from a PDU web UI. Specify it only when that job is real. ## When managed is justified - You need to see kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. - Multi-tenant or billed halls. - The board or BMS does not already meter at this grain. ## When you do not need managed (or do not want outlet switching) - GPU and compute racks that are dual-corded. You do not power-cycle a training node from a PDU web UI. - EPMS or BMS already meters the feed. A metered PDU can be redundant. - The PDU management network is another plane to design, secure, and cable. Do not add it by default. - Unmanaged, or metered but not switched, is a valid spec. Cheaper, fewer failure modes, still distributes power. Pages: /what-we-build#pdu and /capabilities#pdu. Power plant layer. Not a SKU list. Canonical: https://neuronplant.com/capabilities#pdu # Facility as architecture of the hall we occupy The hall is architecture. Rack space is not power or cooling. A hall we occupy is a set of coupled limits: U, weight, floor load, normal and peak kW, A/B, PDU, heat, clearance, and pathways. Empty U with a dead feed is not capacity. The hall owner supplies the room. NeuronPlant engineers the GPU / TPU / ASIC cell to fit, or matches another hall. We do not pour the slab. - U / rack space: Necessary and almost never sufficient. Count the accelerators, CDU manifolds, patch, and service U. Then ask whether power and heat still fit. - Weight: Liquid-cooled GPU racks and full storage frames are plant loads. If the slab or raised floor cannot take them, the SKU does not ship, or it ships and sits. - Floor load: kN/m², not a vibe. Point loads under NVL-class racks and CDUs. Path in from the dock matters as much as the pad. - Normal and peak kW: Nameplate is not the bill. Training peaks. Inference is different. Design the feed, UPS, and breaker to the number that shows up, A/B where the architecture requires it. - A/B: Dual cord is a drawing: two PDUs, two paths, maintenance bypass. Dual cord on a single upstream is a sticker. - PDU: Circuit count, phase, and whether managed is justified. See Power / PDU. Reserve circuits for the additive cell or admit the row is full. - Heat: Rejection path, not a CRAH brochure. Air until physics fails. Then water, CDU, and isolation. Peak kW is peak heat. - Clearance: Service, manifolds, and a sled that can come out without tensioning fiber. A packed hot aisle is not density. It is an unmaintainable plant. - Pathways: Fiber baskets, power, liquid. Overhead versus underfloor. The next trunk needs a reserved path. See Cabling / tomorrow. Pages: /capabilities#facility, /what-we-build#site. Expansion gate: /expansion#gates. Canonical: https://neuronplant.com/capabilities#facility ## NVIDIA platforms NVIDIA DGX B200 / NVIDIA DGX B300 nodes and NVIDIA DGX GB200 NVL72 / NVIDIA DGX GB300 NVL72 racks, cited from public NVIDIA pages. HTML: https://neuronplant.com/models Split file: https://neuronplant.com/llms/platform.txt # NVIDIA DGX B200 Class: Node GPUs: 8x NVIDIA Blackwell Unified develop-to-deploy node. NVIDIA lists 1,440 GB GPU memory and up to 400 Gb/s ConnectX-7 / BlueField-3 ports. NVIDIA source: https://www.nvidia.com/en-us/data-center/dgx-b200/ NeuronPlant does not claim these as original measurements. Specs are NVIDIA's. The room, loop and fabric are the design. These size a GPU plant. Compute is accelerators: GPU, TPU, and other ASICs. We do not maintain a weekly chip catalog. Canonical: https://neuronplant.com/models # NVIDIA DGX B300 Class: Node GPUs: 8x NVIDIA Blackwell Ultra NVIDIA positions NVIDIA DGX B300 for reasoning / inference density. Docs list 8x B300 SXM, ConnectX-8 up to 800 Gb/s, BlueField-3 for storage and management. NVIDIA source: https://www.nvidia.com/en-us/data-center/dgx-b300/ NeuronPlant does not claim these as original measurements. Specs are NVIDIA's. The room, loop and fabric are the design. These size a GPU plant. Compute is accelerators: GPU, TPU, and other ASICs. We do not maintain a weekly chip catalog. Canonical: https://neuronplant.com/models # NVIDIA DGX GB200 NVL72 Class: Rack GPUs: 72 Blackwell + 36 Grace Liquid-cooled rack-scale NVLink domain. NVIDIA describes the GB200 NVL72 architecture as a 72-GPU NVLink domain for large-model inference and MoE. NVIDIA source: https://www.nvidia.com/en-us/data-center/gb200-nvl72/ NeuronPlant does not claim these as original measurements. Specs are NVIDIA's. The room, loop and fabric are the design. These size a GPU plant. Compute is accelerators: GPU, TPU, and other ASICs. We do not maintain a weekly chip catalog. Canonical: https://neuronplant.com/models # NVIDIA DGX GB300 NVL72 Class: Rack GPUs: 72 Blackwell Ultra + 36 Grace NVIDIA DGX rack-scale system. Published networking includes ConnectX-8 at 800 Gb/s InfiniBand and NVLink Switch System. NVIDIA source: https://www.nvidia.com/en-us/data-center/dgx-gb300/ NeuronPlant does not claim these as original measurements. Specs are NVIDIA's. The room, loop and fabric are the design. These size a GPU plant. Compute is accelerators: GPU, TPU, and other ASICs. We do not maintain a weekly chip catalog. Canonical: https://neuronplant.com/models ## NVIDIA switches Quantum InfiniBand and Spectrum-X Ethernet families, then cited SuperPOD names for those jobs. HTML: https://neuronplant.com/capabilities#nvidia-switches Split file: https://neuronplant.com/llms/switch.txt # Quantum-X800 Q3400 Fabric: East-West InfiniBand compute Role: Scale-out training / NVL scale-out. East-West compute fabric. NVIDIA SuperPOD B300 RA: 800 Gb/s, rail-optimized fat tree for the compute fabric. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/dgx-superpod-components.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Quantum-2 QM9700 Fabric: InfiniBand compute or storage/data Role: NDR 400 Gb/s leaf. Compute fabric or dedicated storage/data fabric, depending on the SuperPOD job. NVIDIA SuperPOD RA uses MQM9700-NS2R (AC) for InfiniBand storage. Quantum-2 is 64x 400 Gb/s NDR in 1U. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/network-fabrics.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Spectrum-4 SN5600 Fabric: Spectrum-X Ethernet Role: AI Ethernet. Plane depends on the job: East-West compute, storage/data, or North-South front-end. Do not assume one role. 64x 800GbE, Spectrum-4 ASIC. NVIDIA SuperPOD Ethernet RA uses SN5600 / SN5610 for storage and in-band paths. NVIDIA source: https://www.nvidia.com/en-us/networking/spectrumx/ Canonical: https://neuronplant.com/capabilities#nvidia-switches # Spectrum SN5610 Fabric: North-South / user Role: User storage, NFS, homes, application-facing path. Front-end / North-South, not the high-performance storage/data fabric. NVIDIA SuperPOD B300 RA connects user storage to leaf-layer SN5610 (typically 100G-class DR1 for home / metadata). NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/network-fabrics.html Canonical: https://neuronplant.com/capabilities#nvidia-switches # Four fabric planes Four planes, not four equal cable plants of the same job. East-West compute fabric is GPU-to-GPU, DGX-to-DGX, scale-out: RDMA/RoCE or InfiniBand, NVIDIA Spectrum-X or Quantum-X. Storage/data is its own internal data plane when physically separate. Do not fold it into a generic East-West line if dedicated. North-South is users, apps, API gateway/LB, enterprise/core. Not another GPU rail. Management/OOB is BMC and operational control. Never application North-South. Converged versus dedicated is a decision on whether compute and storage share a leaf. NVIDIA SuperPOD public RAs already split compute, storage, and user/NFS switch jobs. Structured cabling builds those plants. It is not a fifth fabric. We do not sell a weekly RoCE tune as a service SKU. Congestion behaviour is a plant job: lossless where it is justified, QoS, PFC and ECN as mechanisms, not as a product name. ## Hierarchy EAST-WEST: Compute Fabric; Storage/Data Fabric where dedicated. NORTH-SOUTH: Front-End / Application / Enterprise Fabric. MANAGEMENT: OOB / BMC / Management Fabric. Structured cabling is how those plants are built, not a fifth fabric. ## Topologies - Dedicated planes. When: Collectives and checkpoint burst must not share queues. Training halls, SuperPOD-class jobs, noisy neighbours. Storage/data stays off the East-West compute rails. Cost: More leaves, more trunks, more trays. Cleaner failure domains. Easier to expand one plane without recabling the other. - Converged fabric. When: A smaller cell, a bounded RAG/inference job, or a hall that cannot take a second leaf class. Still name oversub and QoS in writing. Convergence is a compute/storage decision, not permission to mix North-South or OOB onto GPU rails. Cost: Fewer SKUs. Shared congestion and a shared blast radius. The first transition is often splitting the plane you collapsed to save ports. ## Planes to name - East-West / East-West Compute Fabric: Collectives, NCCL, scale-out between nodes or NVL racks. The model runtime lives here. This is not user traffic and not BMC. - East-West / Storage / Data Fabric: Compute to high-performance storage: datasets, checkpoints, RAG index path. Do not fold it into a generic East-West compute line if dedicated. Do not auto-label it North-South or East-West without that context. Decide with the storage machine, not with leftover ports. - North-South / North-South / Front-End Fabric: Application-facing Ethernet. NFS homes, logs, ingest, people. Must not share queues with training I/O. Not just another rail next to the GPU rails. - Management / Management / OOB Fabric: Lights-out, firmware, serial, leak, PDU management if that job exists. In-band plant services sit here, not on application North-South. If this dies inside a compute bundle, the hall is dark and you cannot see why. ## Networking is more than port count - Topology: Leaf/spine, rail alignment, how a new leaf joins. Not a pile of ToR switches. - East-West compute: GPU-to-GPU and DGX-to-DGX diameter, RDMA/RoCE or InfiniBand, Spectrum-X or Quantum-X. The model runtime lives here. Not a generic line that also swallows storage. - Storage / data path: Whether checkpoints and RAG share the compute rails. When dedicated, this is its own internal data plane. Do not auto-label it North-South or East-West without that context. - North-South / front-end: Users, applications, API gateway/LB, enterprise/core. Application-facing Ethernet. Not a GPU rail. - Management / OOB: BMC and operational control. Never classify as application North-South. - Oversubscription: Written per plane. A number that looked fine at eight nodes can fail at the first additive cell. - RoCE / lossless Ethernet: A job when Ethernet carries RDMA on the East-West compute fabric. Design congestion, not a slogan. We do not productise a PFC/ECN recipe as a SKU. - QoS, PFC, ECN: Mechanisms that keep planes from eating each other. Named on the drawing. Tuned in commissioning. Not a catalog line. - Optics: Reach, power, and what the structured cabling can take. 800G on dirty MPO is inventory. - Failure domains: What dies with one leaf, one trunk, one PDU. Blast radius is architecture. - Expansion sequence: Which leaf, which trunk, which reserved position. See /expansion. Port count is not a sequence. Pages: /capabilities#fabrics. Related: /cabling, /expansion, /insights/ethernet-vs-infiniband. Canonical: https://neuronplant.com/capabilities#fabrics ## Storage by workload Training, RAG, inference, and HPC mixes as different machines. Large inference also has a KV / prefix cache tier. HTML: https://neuronplant.com/capabilities#storage Split file: https://neuronplant.com/llms/storage.txt # Storage / Training / post-training Look for: Sustained sequential GB/s and checkpoint burst, not IOPS marketing. Shared parallel filesystem for datasets. Separate checkpoint target that can absorb a simultaneous flush from the job. NVIDIA SuperPOD B300 RA states high-performance storage I/O per node must exceed 40 GB/s (Ethernet RA) and, on the Quantum-X800 RA, 80 GB/s. Typically: HPS on IB (QM9700) or RoCE (SN5600) as the storage/data fabric, plus a quieter NFS/user tier on SN5610 as North-South / front-end. Do not fold HPS into a generic East-West compute line if dedicated. Avoid: A single enterprise NAS asked to be both home directories and the training scratch. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/network-fabrics.html Canonical: https://neuronplant.com/capabilities#storage # Storage / RAG / enterprise retrieval Look for: Random read IOPS, metadata, and a clean split between documents and the vector index. Object or file store for source documents. NVMe-backed or in-memory vector index. Ingest path sized to nightly rebuilds, not to training checkpoints. Typically: Object + vector engine. North-South / front-end Ethernet is usually enough if the index is local or on a low-latency tier. That is not the East-West compute fabric. Avoid: Putting the vector store on the same parallel FS as a training job. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/capabilities#storage # Storage / Inference / NIM serving Look for: Fast model-weight load, local NVMe or a warm cache, isolation from training I/O, and a KV / prefix cache plan on large serving. NVIDIA NIM selects a profile from its support matrix (precision, tensor-parallel size, LoRA). Weights load from local NVMe. That is not the whole inference memory story. On a large workload, KV cache needs a hierarchy: GPU HBM, host overflow, and sometimes a shared prefix cache across the serving plane. That tier is not a storage array and not a checkpoint pool. Typically: Node-local NVMe for NIM images and weights. KV in HBM first. Host memory and a serving-plane prefix cache when context and concurrency fill the GPU. Avoid: Loading 70B-class weights over a contended NFS home share. Treating KV as another LUN on the training filesystem. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/capabilities#storage # Storage / Research / HPC mix Look for: Scratch that can be purged, project space that cannot, and a scheduler that knows the difference. Burst scratch on the HPS fabric. Project datasets on a capacity tier. Home and logs on the user-storage Ethernet path, as in the NVIDIA SuperPOD split. Typically: Two storage systems: HPS + user storage. NVIDIA documents this split explicitly. Avoid: One quota, one protocol, every lab on the same queue. NVIDIA source: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/dgx-superpod-components.html Canonical: https://neuronplant.com/capabilities#storage # Storage sizing Capacity is not a storage architecture. TB on a quote is not a spec. Engineer the requirement before the product: throughput, IOPS, metadata, checkpoints, RAG, and protection. Then pick a machine that can do that job. Training, RAG, and inference remain different machines. - Throughput: Sequential GB/s for training scratch and checkpoint burst. NVIDIA SuperPOD public RAs already publish per-node floors. Size to the job, not to a single array's headline. - IOPS: Random read for RAG retrieve and metadata-heavy ingest. A parallel FS built for sequential train I/O is the wrong machine. - Metadata: Small files, many objects, index catalogs. The namespace can saturate before the bytes do. - Checkpoints: A simultaneous flush is a storage event. Burst target and recovery window are part of the plant, not an afterthought quota. - RAG: Source corpus, embeddings, and the index are different sizes and different I/O. See RAG storage below. Do not size from vector capacity alone. - Protection: Copies, erasure, snapshots, rebuild time. Protection is throughput you no longer have for the job while a disk is gone. ## RAG storage is sized from source data, not from vector capacity. Start with raw TB and how it grows. Then decide copy versus reference, retention, embedding volume, protection, and what a re-index costs. The vector store is one downstream number. It is not the architecture. - Raw TB: Authoritative source size today. Everything else is derived. - Growth: Ingest rate and how long the plant must take it without a storage transition. - Copy vs reference: Landing zones that copy the corpus multiply TB. Reference in place saves bytes and couples failure domains. Write the choice. - Retention: How long each data state lives. Legal hold is not the same as hot retrieve. - Embeddings: A function of chunking, dimensions, and versions kept. A new embedding model is a storage event, not a config flag. - Protection: Which states are rebuilt from source, which are backed up, which are disposable. - Re-index: Bytes, IOPS, and window to rebuild the index from chunks. If you cannot re-index, you do not own the RAG plant. Pages: /capabilities#storage, /capabilities#rag-storage. RAG path and lifecycle: /applications#rag. Canonical: https://neuronplant.com/capabilities#sizing ## DSX campus How we read public NVIDIA DSX campus layers for a hall partner or customer program. Not a campus NeuronPlant builds. NVOnline packs named only, not reprinted. HTML: https://neuronplant.com/dsx Split file: https://neuronplant.com/llms/dsx.txt # DSX campus layer: Grid / substation Utility lands above 100 kV, steps into a 34.5 kV campus backbone with two buses: Network Core and GPU Compute. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: BESS and standby generation BESS on the 34.5 kV bus as a buffer for spiky AI load. Standby generation in N+2 for Core continuity. Sizing is site-specific. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: Central Utility Building Shared plant for liquid and air. Chillers on 4.16 kV. Pumps, dry coolers and controls on 480 V. UPS only on the Core side of the CUB. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: Dry coolers Air-side heat rejection, no evaporative water in this reference pattern. NVIDIA states a 45°C liquid-cooling design point so more of the facility budget stays with compute. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: CDU gallery Liquid-to-liquid CDUs in the mechanical gallery, N+1 groups. Secondary loop to GPU cold plates. NVIDIA publishes a TCS design flow of at least 1.5 LPM/kW. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: Compute data hall A Scalable Unit is one compute hot-aisle plus one support hot-aisle. Cabinet TDP in the public overview runs from 198 kW (MGX Gen 1.1) to 330 kW (Vera Rubin NVL72). Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # DSX campus layer: CIN spine and network core East-West cluster interconnect with a 500 m optical reach limit. That number is a fiber plant constraint (SMF, MPO, polarity, test), not a switch SKU. Network Core holds tenant access (North-South), secure management (not application North-South), high-speed storage (its own data fabric when dedicated) and meet-me rooms. Restated from NVIDIA's public DSX Facilities Infrastructure Reference Design Overview. NeuronPlant is not an NVIDIA Cloud Partner and does not reprint NVOnline packs. Canonical: https://neuronplant.com/dsx # Named NVIDIA DSX packs Titles and document numbers only. No drawings, no BOM, no copied pages. - NVIDIA Vera Rubin NVL72 Reference Design (NVOnline #1151654, Compute / cluster): Generation-specific AI factory pack for Vera Rubin NVL72. NVIDIA lists it on the public DSX hub as NVOnline-gated. - NVIDIA DSX Vera Rubin Facilities Infrastructure Reference Design (NVOnline #1145739, Site / facilities): Facilities pack for the Vera Rubin generation. Same NVOnline gate. We do not reproduce the drawings or BOM. - NVIDIA DSX Facilities Infrastructure Design Guide (NVOnline #1152370, Campus design guide): NVIDIA's public overview cites Design Guide v2.0 (19 August 2026) as the rendering source. The full guide is NVOnline. Canonical: https://neuronplant.com/dsx # NVIDIA published example, not a NeuronPlant project 250 MW-class IT load / 96 Scalable Units - Site footprint: 157 acres - Building: 972,500 SF - Compute halls: 4 - GPU racks: 1,536 - GPUs: 110,592 - Compute IT + mechanical: 240 MW - Core IT + mechanical: 22.8 MW NVIDIA presents this table as a sizing reference independent of the campus rendering. It is not a NeuronPlant project, not a campus we construct, and not a hall we sell. Actual sites differ. GPU rack count in that example assumes a maximum of eight racks per compute row. Canonical: https://neuronplant.com/dsx ## Factory templates Private AI, training, AI-as-a-Service, sovereign, and large-scale factories NeuronPlant can deliver. HTML: https://neuronplant.com/ai-factories Split file: https://neuronplant.com/llms/factory.txt # Enterprise Private AI Factory Private enterprise AI / RAG / inference. A controlled environment for enterprise models, retrieval, and inference, inside the organization's security and data boundary. ## Pipeline - Data Sources - Ingestion - Object / File Storage - Vector / RAG Layer - Compute - Inference Platform - Enterprise Applications - Operate ## Infrastructure - Site: Inside the organization's security and data boundary. - Power: Rack-level redundancy, typically 20-40 kW/rack - Cooling: Air or targeted liquid depending on density - Compute: Accelerators for NIM / RAG serving. GPU plant: NVIDIA DGX B200 or NVIDIA DGX B300, or HGX. TPU and other ASICs when that is the job. - Network: Spectrum-X East-West compute fabric, dedicated storage/data path, separate North-South front-end - Structured cabling: Labeled LC/MPO plant. Management/OOB on its own patch field. Factory-tested trunks. Polarity map before first optic. - Storage: Object for documents + NVMe or memory vector index. Model weights on local NVMe. - Platform: Sources → ingest → RAG / inference → enterprise applications. - Operate: Docker / Kubernetes cell. Observe as part of the plane. ## Considerations - Dataset size and retention - Growth of concurrent sessions - Availability target - Security and data residency - Latency to enterprise applications - Data lifecycle and audit - Whether the existing hall plant can take 400G without a recable - How the serving plane is observed after handover Canonical: https://neuronplant.com/ai-factories/private-ai # High-Performance Training Factory High-performance accelerator training infrastructure. A tightly coupled training cluster: fabric, storage bandwidth, and cooling sized to the job, not the idle rack. GPU is the common case. TPU and other ASICs when that is the domain. ## Pipeline - Datasets - Staging / Parallel FS - Training fabric - Checkpoint Store - Scheduler - Experiment Platform - Operate ## Infrastructure - Site: Dedicated training hall we occupy. Floor, water, and power are gates of that room before the first node. - Power: High-density feeds, often 40-120 kW/rack - Cooling: Direct liquid cooling and CDU-backed rejection - Compute: Accelerators in one training domain. GPU plant: NVIDIA DGX B200 or NVIDIA DGX B300 nodes, or NVIDIA DGX GB200 NVL72 or NVIDIA DGX GB300 NVL72 racks. TPU and other ASICs sized the same way, not as a catalog. - Network: Quantum InfiniBand East-West compute fabric; dedicated storage/data fabric. Cited SuperPOD jobs on Capabilities. - Structured cabling: Rail-aligned MPO trunks. Type-B polarity as a drawing, not a field guess. East-West compute, storage/data, and management/OOB on separate trays. IL and polarity records before first job. - Storage: HPS parallel FS sized to NVIDIA SuperPOD guidance (>40 GB/s per NVIDIA DGX B300 node; >80 GB/s on XDR RA) plus a checkpoint burst target - Platform: Datasets → staging → train fabric → checkpoints → scheduler. - Operate: 3+3 Kubernetes cell. Observe the fabric and the job. ## Considerations - Model scale and parallelism strategy - Job mix vs dedicated cluster - Checkpoint frequency and recovery - Fabric congestion under all-reduce - Structured cabling: polarity, slack, and whether the hall can be tested - Facility power and water lead times - Expansion without re-architecture - Who watches the fabric and the job after first run Canonical: https://neuronplant.com/ai-factories/training # Multi-Tenant Compute Factory Multi-tenant accelerator infrastructure designed for service providers. A factory built to sell capacity: isolation, metering, rapid provisioning, and a hall-partner model that can grow in cells. ## Pipeline - Onboarding - Tenancy / Isolation - Accelerator pools - Shared Fabric - Platform Control Plane - Billing / Observability - Operate ## Infrastructure - Site: Colo or owned hall from a hall partner, grown in cells. Not a NeuronPlant campus. - Power: Cell-based capacity, N+1 at hall and rack - Cooling: Liquid-ready rows for density growth - Compute: Pooled accelerators. GPU nodes with MIG / tenant slicing where applicable. - Network: Segmented Spectrum-X or Quantum fabrics: East-West compute isolated from tenant North-South, noisy-neighbor controls - Structured cabling: Per-plane trays so a tenant storage change cannot yank an East-West compute trunk. Labels that survive a night shift. - Storage: Per-tenant namespaces on a shared high-performance backend - Platform: Onboarding → tenancy → accelerator pools → control plane → customer workloads. - Operate: Shared control plane, per-tenant isolation, observe as a billed plane. ## Considerations - Tenant isolation model - Oversubscription policy - Support and SLA boundaries - Chargeback and telemetry - Cable plant isolation between tenants and between planes - Cooling headroom for densification - Sovereign or regional cells - Tenant telemetry without a dashboard shopping list Canonical: https://neuronplant.com/ai-factories/ai-as-a-service # Sovereign AI Factory Private, controlled AI infrastructure for regulated organizations. Infrastructure that can be operated, inspected, and expanded under a defined jurisdiction: power to platform, not only the model weights. ## Pipeline - National / Org Data - Controlled Ingest - Sovereign Storage - Private compute domain - Governed Platform - Authorized Workloads - Operate ## Infrastructure - Site: In-jurisdiction hall from a partner or custom-system supplier. Inspectable plant and operational control. - Power: Facility-level independence and documented failover - Cooling: Heat rejection the hall owner or custom-system supplier presents, with operational control in-jurisdiction. - Compute: Dedicated accelerator domains, no shared cloud control plane - Network: Air-gapped or tightly gated interconnects - Structured cabling: Inspectable plant: labeled, tested, as-built. No undocumented jumpers on the air-gap. - Storage: In-country datasets, encryption and key custody - Platform: National / org data → controlled ingest → governed platform → authorized workloads. - Operate: In-jurisdiction runtime cell. Observe stays inside the boundary. ## Considerations - Data residency and custody - Personnel and operational control - Supply chain of compute - Auditability of the stack - Independent support model - Long-term expansion path - Observe stays inside the jurisdiction Canonical: https://neuronplant.com/ai-factories/sovereign-ai # Large-Scale AI Factory Large-scale accelerator environment with liquid cooling and high-speed networking. A hall-scale accelerator program delivered with a hall partner or custom-system supplier: megawatts and liquid as hall constraints, 800G fabrics and the GPU / TPU / ASIC plant as NeuronPlant's job. Not NeuronPlant as the data-center contractor. ## Pipeline - Datasets / ingest - Hall-scale HPS - Accelerator supercluster - Checkpoint / archive - AI platform - Production workloads - Operate ## Infrastructure - Site: MW-class hall from a partner or custom-system supplier. Utility and water are gates of that room, not a NeuronPlant civil works package. - Power: MW-class, staged to utility and on-site distribution - Cooling: Facility water, CDUs, rear-door or direct-to-chip, N+1 rejection - Compute: Accelerator halls at multi-pod scale. GPU plant: NVIDIA DGX GB200 NVL72 / NVIDIA DGX GB300 NVL72. TPU and other ASICs when that is the job. - Network: Quantum-X or Spectrum-X East-West scale-out, dedicated storage/data fabric, North-South for users, management/OOB independent - Structured cabling: SMF/MPO campus plant inside the CIN 500 m class from the public NVIDIA DSX overview. Overhead baskets. Trays grouped by plane: East-West compute, storage/data, North-South, management/OOB. OTDR on backbone. As-builts before first power-on. - Storage: Multi-tier: scratch, dataset, checkpoint, archive - Platform: Datasets → hall HPS → supercluster → checkpoints → production workloads. - Operate: Cell-by-cell Kubernetes. Observe before first training run, not after. ## Considerations - Utility interconnection timeline (hall owner or custom-system supplier) - Water and heat-rejection path - Phased pod delivery - Logistics of high-density racks - Fiber plant reach, polarity, and test records at hall scale - Operations before first training run - Cell-by-cell expansion - Observe before first training run Canonical: https://neuronplant.com/ai-factories/large-scale ## Design decisions Why this fabric, cooling, storage, sequence, or plant. Rationale in writing. HTML: https://neuronplant.com/mcp Split file: https://neuronplant.com/llms/decision.txt # Four fabric planes, not four equal jobs Topic: fabric Kind: engineering-position ## Question Why not one flat network for the whole AI factory? ## Decision Design four planes. East-West compute fabric (GPU-to-GPU, DGX-to-DGX, RDMA/RoCE or InfiniBand, Spectrum-X or Quantum-X). Storage/data fabric as its own internal data plane when dedicated. North-South / front-end for users, apps, API gateway/LB, enterprise/core. Management/OOB for BMC and operational control, never as application North-South. They may share a vendor. They do not share a failure domain. ## Why Collectives, checkpoint burst, user traffic, and BMC traffic have different congestion and blast-radius profiles. Collapsing them to save ports is how an outage looks like a model problem. NVIDIA SuperPOD public reference architectures already name distinct switch jobs for those planes (Q3400 or QM9700 compute, QM9700 or SN5600 storage, SN5610 user/NFS). Structured cabling is how those plants are built, not a fifth fabric. ## Alternatives considered - Single Ethernet leaf for every job - InfiniBand for every job including homes and BMC ## Related - https://neuronplant.com/insights/ethernet-vs-infiniband - https://neuronplant.com/capabilities#nvidia-switches - https://neuronplant.com/capabilities#fabrics Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Specify the loop before the NVL72 purchase order Topic: cooling Kind: engineering-position ## Question Why is an NVL72 not treated as a dense server install? ## Decision Treat NVL72 as plant equipment. Facility water, CDU, isolation, and a 100 kW-class feed are gates. The rack is specified after the room can take it. ## Why An NVL72 is a liquid-cooled NVLink domain. If the hall cannot present the water, the rack ships and sits. Redundancy lives in the loop, the CDU, and the rejection plant, not in a spare fan. Power nameplate and sustained draw are different numbers. Design to the one that shows up on the utility bill and the breaker. ## Alternatives considered - Order the rack and retrofit cooling later - Air-cool a 72-GPU NVLink domain as if it were 8-GPU nodes ## Related - https://neuronplant.com/insights/deploy-nvl72 - https://neuronplant.com/dsx Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Storage is sized to the workload, not to one NAS brochure Topic: storage Kind: engineering-position ## Question Why not one enterprise array for training, RAG, and inference? ## Decision Split the machines. Training wants sequential GB/s and checkpoint burst. RAG wants random read IOPS and a document / index split. Inference wants local NIM weights. Homes and logs stay on a quieter user tier. ## Why NVIDIA SuperPOD B300 public guidance states high-performance storage I/O per node must exceed 40 GB/s (Ethernet RA) and 80 GB/s on the Quantum-X800 RA. A single NAS asked to be both home directories and training scratch will stall jobs and look like a GPU problem. The index does not belong on the training filesystem. ## Alternatives considered - One quota, one protocol, every lab on the same queue - Vector store on the training parallel FS ## Related - https://neuronplant.com/capabilities#storage - https://neuronplant.com/capabilities#rag-storage - https://neuronplant.com/models Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Feasibility before architecture before procurement Topic: power Kind: engineering-position ## Question Why not start from the accelerator SKU list? ## Decision Sequence is feasibility first (what the hall can present: power, water, space, floor, connectivity), architecture of the accelerator cell second, procurement third. ## Why Time-to-power and time-to-water bind the cell more often than PUE slogans. Those clocks belong to the hall owner or the custom-system supplier. If you cannot draw the path the room presents to the compute rack, including maintenance bypass, you do not yet have an architecture. NeuronPlant does not become the utility or the civil contractor. ## Alternatives considered - Buy accelerators, then discover the hall cannot present the feed - Redundancy that exists only on a slide - NeuronPlant pouring a campus to rescue a bad sequence ## Related - https://neuronplant.com/insights/power-cooling-ai-factories - https://neuronplant.com/what-we-build Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Specify the PDU job. Managed is not automatic. Topic: power Kind: engineering-position ## Question Why not put a managed PDU in every compute rack? ## Decision Treat unmanaged distribution, metered, and switched/outlet-level as three jobs. Write the one you need. Do not default to a management plane. ## Why A PDU is how the rack gets power. Managed means metering, a network, and often remote outlet control. That is justified when you need to see kW, phase balance, and headroom versus nameplate, when the hall is multi-tenant or billed, or when the board and BMS do not already meter at that grain. Dual-cord GPU and compute racks do not want outlet switching. You do not power-cycle a training node from a PDU web UI. If EPMS or BMS already meters the feed, a metered PDU can be redundant. The PDU management network is another plane to design, secure, and cable. Unmanaged, or metered but not switched, is a valid spec. Cheaper, fewer failure modes, still distributes power. ## Alternatives considered - Managed PDU in every rack as a standard BOM line - Outlet switching on dual-cord training nodes - Metered PDUs on top of an EPMS or BMS that already meters the feed ## Related - https://neuronplant.com/capabilities#pdu - https://neuronplant.com/what-we-build#pdu - https://neuronplant.com/insights/power-cooling-ai-factories Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Specify structured cabling with the fabric, not after the GPUs land Topic: cabling Kind: engineering-position ## Question Why is structured cabling and passive infrastructure a factory layer instead of a fit-out? ## Decision Four fabric planes on separate trays (East-West compute, storage/data, North-South front-end, management/OOB), polarity as a drawing, IL / OTDR / as-builts as gates. Recable is the recovery when this is skipped. Firmware is not. Structured cabling is how those plants are built, not a fifth fabric. ## Why Dirty MPO, mixed polarity, no slack, blocked chimneys, and mixed trays take rails dark while teams chase NCCL and switch code. NVIDIA's public DSX overview puts a 500 m optical reach limit on the CIN spine. That is a fiber plant number. NeuronPlant has watched an unnamed operator buy a large CPU and GPU fleet, hand structured cabling to another firm, and sit offline. That is a cautionary record, not a customer logo. ## Alternatives considered - Let cabling to a fit-out contractor after hardware lands - Nearest free port installs on mixed trays ## Related - https://neuronplant.com/cabling - https://neuronplant.com/cabling#tomorrow - https://neuronplant.com/insights/passive-plant-offline - https://neuronplant.com/projects/gpu-hall-passive-plant Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Use public NVIDIA DSX as language for a hall partner or customer program, not as a campus NeuronPlant builds Topic: campus Kind: engineering-position ## Question Why cite NVIDIA DSX if NeuronPlant is not an NVIDIA Cloud Partner and does not construct data centers? ## Decision Restate the public facilities overview when a hall partner or customer program needs that language. Name NVOnline packs by title and number only. Do not reprint them. Do not claim partnership. Do not present NVIDIA's 157-acre example as a NeuronPlant delivery. ## Why The open DSX campus (grid, BESS, CUB, dry coolers, CDU gallery, compute hall, CIN) is the right shape of the problem at hall and campus scale for the people who actually own the room. The gated packs are NVIDIA's. Shipping them, or implying NVOnline access we do not have, is not engineering. It is a compliance failure. Treating the 157-acre table as a NeuronPlant campus is a positioning failure. ## Alternatives considered - Treat a brochure rack as a campus design - Copy NVOnline drawings into a customer pack - Sell NeuronPlant as the builder of NVIDIA's example campus ## Related - https://neuronplant.com/dsx Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Accelerators without a data-to-workload path are inventory Topic: pipeline Kind: engineering-position ## Question Why is the AI platform a plant layer? ## Decision Architect ingest, storage tiers, train or RAG, runtime, NIM serving from public NVIDIA docs, and handover as part of the factory. Do not claim to write every model or to be a generic software ISV. ## Why After power, cooling, racks and accelerators exist, bytes still have to move. Train / fine-tune and RAG are different I/O machines that can share a hall. Mixing them on one filesystem is how jobs stall. Ownership of the application stays with the operator. ## Alternatives considered - Stand up accelerators and call the factory done at first token - Bundle a NeuronPlant SaaS application suite ## Related - https://neuronplant.com/applications - https://neuronplant.com/applications#rag - https://neuronplant.com/runtime - https://neuronplant.com/models Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Run the factory as a Docker and Kubernetes cell, not as ssh-and-excel Topic: pipeline Kind: engineering-position ## Question Why Docker and Kubernetes on an AI factory, and why three masters? ## Decision Ship serving, ingest, RAG, NIM runtime, and plant services as Docker images. Schedule them on Kubernetes. Start with three masters and three workers. Masters may be VMs. Workers that hold accelerators are usually bare metal. Fluent Bit is the example log shipper. Do not pin CNI, mesh, or operator SKUs. ## Why The hall is not a unique snowflake OS. Same artifact from lab to factory. An AI factory is a fleet. A fleet that is ssh-and-excel will not stay up. Three masters so etcd and the API stay up if one dies. That is why 3, not 1. Observability is part of the plane, not a dashboard catalog. ## Alternatives considered - Treat each hall as a unique OS and copy binaries by hand - Run production on a single master - Require three physical master servers when VMs will hold quorum - Maintain a weekly CNI / mesh / accelerator-operator SKU list ## Related - https://neuronplant.com/runtime - https://neuronplant.com/runtime#plane - https://neuronplant.com/applications#runtime Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp # Growth is an architecture evolution, not a node count Topic: expansion Kind: engineering-position ## Question Why not sell scale-from-two-to-eight-DGX as the growth story? ## Decision Name additive expansion versus an architecture transition. Write how far Day-1 grows without a redesign. Gate the next investment on compute, storage, fabric, platform, and facility measurements. Offer minimum Day-1 or a future-ready foundation, with the money path visible. Do not promise unlimited scale. ## Why Adding nodes inside a designed envelope is one job. Changing fabric class, cooling, storage machine, or hall is another. A cheaper first invoice that forces a recable is a steered money path. Integrity is showing CapEx versus disruption before procurement. ## Alternatives considered - Scale from 2 to 8 DGX as a slogan - Unlimited scale as an intake option - Hide the future-ready premium so the minimum cell wins the bid ## Related - https://neuronplant.com/expansion - https://neuronplant.com/start-a-project - https://neuronplant.com/cabling#tomorrow Source: https://neuronplant.com/mcp Canonical: https://neuronplant.com/mcp ## Plant notes Short technical notes on NVL72, power/cooling, fabrics, structured cabling, and steered RFPs. HTML: https://neuronplant.com/insights Split file: https://neuronplant.com/llms/insight.txt # What Does It Actually Take to Deploy an NVL72? It is not a server install. It is a power, water, floor, and fabric problem that happens to include 72 GPUs. An NVL72 rack is a liquid-cooled NVLink domain. Treat it as plant equipment, not as a dense 42U server. The first constraint is almost never the purchase order. It is facility water, rack power in the 100 kW class, floor loading, and a network design that does not starve the NVLink domain the moment training starts. Power: plan for a dedicated high-density feed the hall can present, A/B where the architecture requires it, and a PDU/busway story that can be maintained without taking the pod down. Nameplate and sustained draw are different numbers. Design to the one that shows up on the bill and the breaker. Managed vs unmanaged is a constraint: dual-cord training racks usually do not want outlet switching. Cooling: direct-to-chip liquid, CDU, and a heat-rejection path back to the facility water or dry coolers the hall supplies. If the hall cannot present the water, the rack does not ship, or it ships and sits. Redundancy lives in the loop, the CDU, and the rejection plant, not in a spare fan. Network: the scale-up domain is NVLink. The scale-out domain is still a fabric you have to design: 800G class East-West compute, with storage/data and management/OOB as separate planes, and North-South for users and apps. Cabling, optics, and cable management are engineering, not day-two tidy-up. Operations: you need a runbook before first power-on. Leak detection, isolation valves, firmware, BMC, and a commissioning sequence that proves cooling before it proves FLOPS. NeuronPlant's position is simple. Specify the rack only after the room we occupy, the loop, and the feed can take it. We engineer the cell to fit. We do not become the hall. Canonical: https://neuronplant.com/insights/deploy-nvl72 # Power and Cooling Requirements for Modern AI Factories Accelerator density moved AI out of the IT closet and into industrial infrastructure. GPU is the common case. The planning units are kilowatts, liters, and months. Traditional IT racks lived in the 5-15 kW range. AI racks did not stay there. Air-cooled high-density GPU systems already push past what a standard raised floor and CRAH fleet were designed to do. Liquid is not a preference at that point. It is the load path. Power planning starts at what the hall can present, not at the PDU SKU. Available kW, interconnection lead time the hall owner carries, transformer capacity, UPS topology, and the decision between centralized and rack-level redundancy determine the cell more than the SKU list. At the rack, a PDU is how power is distributed. Managed means a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Specify that job when you need to see kW, phase balance, and headroom versus nameplate (training loads are not nameplate), when the hall is multi-tenant or billed, or when the board and BMS do not already meter at that grain. Managed is not automatic. Dual-cord GPU and compute racks do not want outlet switching. You do not power-cycle a training node from a PDU web UI. If EPMS or BMS already meters the feed, a metered PDU can be redundant. The PDU management network is another plane to design, secure, and cable. Unmanaged, or metered but not switched, is a valid spec. A useful rule: if you cannot draw the path the hall presents from its feed to the compute rack, including maintenance bypass, you do not yet have an architecture. Redundancy that exists only on a slide will fail the first time a feed is isolated. Drawing that path is not NeuronPlant building a substation. Cooling planning is a heat-rejection problem. CDUs, facility water temperatures, approach, and what happens when a loop is isolated matter as much as cold-plate selection. Design the failure, then design the steady state. PUE still matters, but for AI factories the binding constraints are often time-to-power and time-to-water: clocks that belong to the hall owner or the custom-system supplier. A perfect cooling plant that arrives a year after the accelerators is not an AI factory. It is inventory. The engineering sequence we use is feasibility first: power, water, space, floor, and connectivity the hall can present. Architecture of the accelerator cell second. Procurement third. That order is what keeps a multi-million-dollar bill of materials from arriving at a room that cannot run it. NeuronPlant does not pour the room to rescue a bad sequence. Canonical: https://neuronplant.com/insights/power-cooling-ai-factories # Ethernet vs InfiniBand for AI Infrastructure The fabric is part of the model runtime. Choose it for collectives, storage, and operations, not for a vendor slogan. Training workloads are sensitive to latency and to congestion during collective operations. Inference and RAG are often more sensitive to isolation, East-West fairness, and the path to storage. One fabric rarely serves every plane well if it is treated as a single flat network. InfiniBand remains a proven choice for tightly coupled training domains. It is not automatically the right choice for every AI factory. Ethernet at 400/800G, with a lossless or properly engineered congestion story, is now a serious training fabric, and often a better fit where the operator already runs Ethernet at scale. NVIDIA SuperPOD reference architectures name Quantum-X800 Q3400 or Quantum-2 QM9700 for compute, QM9700 or Spectrum SN5600 for storage, and SN5610 for user/NFS paths. What actually decides the design: message size and parallelism strategy, number of endpoints, storage protocol, in-house operational skill, and whether the same fabric is expected to carry tenant traffic. We separate planes. East-West compute, storage/data when dedicated, North-South front-end, and management/OOB are different networks even when they share a vendor. Collapsing them to save ports is how you create an outage that looks like a 'model problem'. Management/OOB is never application North-South. There is no poster-child answer. There is a workload, a scale, and an operations team. The architecture should be able to say why InfiniBand, why Ethernet, or why both, in writing, with a diagram, before anyone orders optics. Canonical: https://neuronplant.com/insights/ethernet-vs-infiniband # What Takes a GPU Hall Offline When Structured Cabling Is Wrong The switches can be right. The GPUs can be on the floor. If the structured cabling and passive infrastructure are amateur, the cluster is inventory. People specify GPUs. They specify switches. They treat structured cabling as a fit-out. That is how a hall stays dark after the hardware has arrived. This is not a branding problem. It is bend radius, dirty MPO, mixed polarity, no slack, blocked airflow, an untestable plant, no as-builts, and compute trunks sharing a tray with storage. Any one of those can zero a rail. Together they zero a hall. Bend radius first. An MPO trunk kinked behind a PDU or a CDU manifold will pass a casual look. Under heat and vibration the insertion loss moves. You get CRCs and flaps. The job dies at forty minutes and everyone stares at NCCL. Dirty MPO is the 400/800G classic. One contaminated lane on an MPO-12 or MPO-16 takes the port down. The switch looks guilty. The end-face was never inspected under IEC rules. A click is not a test. Polarity is a drawing. Type A and Type B are not interchangeable. Mix a cassette and a trunk and some lanes light. Some stay dark. Firmware updates do not fix a crossed key. No slack means the first service event is an outage. You cannot slide a GPU sled, reseat a NIC, or swap a QSFP without tensioning fiber. Strain relief belongs at the designed loop, not at the ferrule. Airflow is a cable problem on air-cooled nodes. DAC and AEC are stiff and thick. Pack them into the rear chimney and inlet temperature rises. GPUs throttle, then shut down. Liquid-cooled NVL racks still need a path that does not fight hoses and manifolds. InfiniBand vs Ethernet does not save you if the plants are mixed. Same cage class is not the same fabric. Rail-aligned layout means each GPU rail lands on the matching leaf. A 'nearest free port' install is how collectives look like software. Storage/data, North-South front-end, and management/OOB are separate cable plants from the East-West compute fabric. NVIDIA SuperPOD reference architectures already split those switch planes. The trays have to split too. Pull a storage jumper off a shared basket and you can take training down with it. Kill OOB inside a compute bundle and you cannot even see the BMC. OOB is not application North-South. Campus reach is a fiber number. NVIDIA's public DSX facilities overview puts a 500 m optical reach limit on the cluster interconnect spine. Hall placement and SMF plant have to land inside that class. A 700 m run of leftover MMF is not a CIN. Test is the gate. End-face inspection. Insertion loss per lane, not a trunk average. Polarity check before optics. OTDR on SMF backbone. As-builts tied to the same IDs as the labels. If you cannot isolate a rail at 03:00 from the drawing, you do not have a plant. You have a pile of patch cords. We have watched an unnamed operator buy a large CPU and GPU fleet, hand structured cabling to another firm, and sit offline for a long time. NeuronPlant did not pull that plant. We will not dress it up as a delivery. Recovery is a specified recable. It is not a software workaround. NeuronPlant's position is simple. Specify the structured cabling and passive infrastructure with the fabric. Not after the GPUs land. Canonical: https://neuronplant.com/insights/passive-plant-offline # A Steered RFP Is Not a Plant Spec A factory tender already aimed at one vendor is a capture. Write the constraints first. Leave the SKU until the plant can take it. An RFP that already names the winner is not a specification. It is a capture dressed as process. The first constraint is the plant: site, power, cooling, compute, network, structured cabling, storage, platform, operate. If those are not written, the SKU list is a shopping cart. Advisors get paid. Some get paid more when a particular BOM ships. That is not a secret. What fails is hiding that money path from the customer. NeuronPlant's position is simple. The customer sees the same constraints, options, and money path we see. We do not steer a plant toward the SKU that pays the advisor. Integrity before margin on a steered BOM. If the two conflict, we lose the margin. We do not lose the spec. No games. No tricks. Transparency is the commercial method, not a footnote. Canonical: https://neuronplant.com/insights/steered-rfp ## Cited models NVIDIA NIM support-matrix profiles used to size nodes vs racks. HTML: https://neuronplant.com/models Split file: https://neuronplant.com/llms/model.txt # Llama 3.3 Nemotron Super 49B v1.5 Use: Enterprise reasoning / chat NVIDIA NIM publishes BF16, FP8 and NVFP4 profiles at TP 1, 2, 4 and 8, with and without LoRA. Fits a single NVIDIA DGX B200 or NVIDIA DGX B300 for many FP8/NVFP4 profiles. TP8 is the full 8-GPU node. Size KV cache and concurrency separately. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/models # Nemotron 3 Super 120B-A12B Use: Larger MoE / reasoning NVIDIA NIM lists BF16, FP8 and NVFP4 across TP 1-8, base and LoRA. Needs more GPU memory than a 49B. On Blackwell Ultra (NVIDIA DGX B300) dense FP4/NVFP4 is the practical serving path. Training or distillation belongs on a coupled fabric, not on the inference node. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/models # Trillion-parameter / MoE class Use: Frontier inference NVIDIA describes GB200 NVL72 as a 72-GPU NVLink domain for real-time trillion-parameter LLM inference and MoE. This is a rack (or many racks), liquid loop, and scale-out fabric problem. Do not size it as nine 8-GPU servers on a 100G ToR. NVIDIA source: https://www.nvidia.com/en-us/data-center/gb200-nvl72/ Canonical: https://neuronplant.com/models ## Structured cabling & passive infrastructure Fiber, MPO/MTP, patch panels, pathways, copper, polarity, test records. HTML: https://neuronplant.com/cabling Split file: https://neuronplant.com/llms/cabling.txt # Structured cabling & passive infrastructure Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. The layer that lands the fabric. Skip it and the GPUs sit dark. ## Four fabric planes (not four equal jobs) EAST-WEST: Compute Fabric, and Storage/Data Fabric where dedicated. NORTH-SOUTH: Front-End / Application / Enterprise Fabric. MANAGEMENT: OOB / BMC / Management Fabric. Structured cabling is how those plants are built, not a fifth fabric. - EAST-WEST / East-West Compute Fabric: MPO/MTP trunks, 400/800G optics, rail-aligned GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray. - EAST-WEST / Storage / Data Fabric: Separate MPO or LC plant onto the storage leaf when dedicated Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails. - NORTH-SOUTH / North-South / Front-End Fabric: LC or lower-count MPO, 100G-class typical Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O. - MANAGEMENT / Management / OOB Fabric: Cat6A or dedicated 1/10G optics. Own patch field. BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark. ## Media - DAC (In-rack / adjacent. Typically sub-3 m at 400/800G.) ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend. Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly. - AEC (Active copper. Longer than DAC, still copper.) In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch. Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail. - AOC (In-row to in-hall. Fiber with the optics already on the ends.) When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut. Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant. - MMF (OM4 / OM5) (In-hall SR-class. MPO/MTP trunks.) Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site. Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested. - SMF (OS2) (Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G.) Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU. Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone. ## Practice - Rail-aligned layouts: Each GPU rail lands on the matching leaf. Cables follow the rail map, not the nearest free QSFP. InfiniBand and Ethernet both fail the same way when rails are crossed: collectives look like a software bug. - InfiniBand vs Ethernet cabling: Same cage class does not mean the same plant. Do not mix IB and Ethernet in one trunk. Do not share cassettes. Label the fabric on both ends before the first transceiver seats. - Underfloor vs overhead: High-density GPU halls usually put liquid and power in the floor or the rear, and put fiber in overhead baskets. Underfloor fiber under a liquid loop is how a drip becomes a dark fabric. Pick one primary path. Document the exception. - Cable management: Bend radius is a spec, not a vibe. Slack loops at the rack and at the distribution frame. Strain relief at the NIC, not at the strain of the next installer. Separate baskets per plant so a storage recable cannot yank a compute trunk. - Labeling: Both ends. Unique ID. Fabric, rack, rail, port, polarity type, length. If the label does not match the as-built, the as-built is a drawing, not a plant. - Test: End-face inspection on every MPO. Insertion loss per lane, not per trunk average. Polarity verification before optics. OTDR on SMF backbone. Fail the lane, not the meeting. No test record means the plant is not done. - As-built documentation: Port map, polarity map, tray route, slack locations, test PDFs tied to the label IDs. The next engineer should be able to isolate a rail at 03:00 without calling the person who pulled it. ## Why amateur cabling takes a GPU hall off the air - Bend radius: A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes. - Dirty MPO: One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected. - Mixed polarity: Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks. - No slack: You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage. - Blocked airflow: DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem. - Untestable plant: No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days. - No as-builts: The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name. - Compute and storage on one tray: A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not. ## CIN reach NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF. ## Ports on the switch are not a growth plan. Cabling is architecture. Trunks, MPO counts, pathways, switch positions, and PDU circuit landings decide whether the next cell is additive or a recable. Pre-build and reserve where the first transition is already visible. Do not pull a Day-1 plant that can only ever be Day-1. - Trunks: Structured trunks with spare fibers, not a jumper for every live port. A spare in the trunk is cheaper than a night in the hall. - MPO counts: Count lanes for the optics class you will grow into, not only the optic seated today. Mixing MPO-12 leftovers into an MPO-16 plant is how the next leaf stays dark. - Pathways: Reserved basket and tray for the next trunk. If the only path is full, additive growth is already a transition. - Switch positions: U and power for the next leaf, with the rail map already drawn. A free ToR slot with no trunk landing is not reserve. - PDUs: Circuit count and landing for the next node or leaf. Cabling the packet path while the power path is exhausted is how a hall dead-ends. - Pre-build / reserve: Justified when Day-1 already names the first transition. Not justified as dark fiber for a fantasy campus. Write the gate. See /expansion. Canonical: https://neuronplant.com/cabling # East-West Compute Fabric Band: east-west MPO/MTP trunks, 400/800G optics, rail-aligned GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray. Canonical: https://neuronplant.com/cabling#plants # Storage / Data Fabric Band: east-west Separate MPO or LC plant onto the storage leaf when dedicated Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails. Canonical: https://neuronplant.com/cabling#plants # North-South / Front-End Fabric Band: north-south LC or lower-count MPO, 100G-class typical Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O. Canonical: https://neuronplant.com/cabling#plants # Management / OOB Fabric Band: management Cat6A or dedicated 1/10G optics. Own patch field. BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark. Canonical: https://neuronplant.com/cabling#plants # DAC In-rack / adjacent. Typically sub-3 m at 400/800G. ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend. Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly. Canonical: https://neuronplant.com/cabling#media # AEC Active copper. Longer than DAC, still copper. In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch. Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail. Canonical: https://neuronplant.com/cabling#media # AOC In-row to in-hall. Fiber with the optics already on the ends. When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut. Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant. Canonical: https://neuronplant.com/cabling#media # MMF (OM4 / OM5) In-hall SR-class. MPO/MTP trunks. Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site. Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested. Canonical: https://neuronplant.com/cabling#media # SMF (OS2) Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G. Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU. Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone. Canonical: https://neuronplant.com/cabling#media # Rail-aligned layouts Each GPU rail lands on the matching leaf. Cables follow the rail map, not the nearest free QSFP. InfiniBand and Ethernet both fail the same way when rails are crossed: collectives look like a software bug. Canonical: https://neuronplant.com/cabling#practice # InfiniBand vs Ethernet cabling Same cage class does not mean the same plant. Do not mix IB and Ethernet in one trunk. Do not share cassettes. Label the fabric on both ends before the first transceiver seats. Canonical: https://neuronplant.com/cabling#practice # Underfloor vs overhead High-density GPU halls usually put liquid and power in the floor or the rear, and put fiber in overhead baskets. Underfloor fiber under a liquid loop is how a drip becomes a dark fabric. Pick one primary path. Document the exception. Canonical: https://neuronplant.com/cabling#practice # Cable management Bend radius is a spec, not a vibe. Slack loops at the rack and at the distribution frame. Strain relief at the NIC, not at the strain of the next installer. Separate baskets per plant so a storage recable cannot yank a compute trunk. Canonical: https://neuronplant.com/cabling#practice # Labeling Both ends. Unique ID. Fabric, rack, rail, port, polarity type, length. If the label does not match the as-built, the as-built is a drawing, not a plant. Canonical: https://neuronplant.com/cabling#practice # Test End-face inspection on every MPO. Insertion loss per lane, not per trunk average. Polarity verification before optics. OTDR on SMF backbone. Fail the lane, not the meeting. No test record means the plant is not done. Canonical: https://neuronplant.com/cabling#practice # As-built documentation Port map, polarity map, tray route, slack locations, test PDFs tied to the label IDs. The next engineer should be able to isolate a rail at 03:00 without calling the person who pulled it. Canonical: https://neuronplant.com/cabling#practice # Bend radius A kinked MPO behind a PDU looks fine until temperature and vibration move it. Then you get CRC, flaps, and a training job that dies at 40 minutes. Canonical: https://neuronplant.com/cabling#failure # Dirty MPO One contaminated lane on an MPO-16 takes down an 800G port. The switch looks guilty. The end-face was never inspected. Canonical: https://neuronplant.com/cabling#failure # Mixed polarity Type A cassette on a Type B trunk. Some lanes light. Some stay dark. Teams chase firmware for weeks. Canonical: https://neuronplant.com/cabling#failure # No slack You cannot slide a GPU sled, a CDU hose, or a NIC without tensioning fiber. The first service event is the first outage. Canonical: https://neuronplant.com/cabling#failure # Blocked airflow DAC and AEC packed into the rear chimney of an air-cooled node. Inlet rises. GPUs throttle, then shut down. The fabric was never the problem. Canonical: https://neuronplant.com/cabling#failure # Untestable plant No labels, no test points, no polarity map. Every incident is a hunt. Mean time to isolate is measured in days. Canonical: https://neuronplant.com/cabling#failure # No as-builts The only record is in someone's head, and that person is not on the ticket. You cannot recable what you cannot name. Canonical: https://neuronplant.com/cabling#failure # Compute and storage on one tray A storage change yanks a compute trunk. NCCL dies. The storage vendor is 'done'. The hall is not. Canonical: https://neuronplant.com/cabling#failure # CIN reach NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF. Canonical: https://neuronplant.com/cabling#media # Ports on the switch are not a growth plan. Cabling is architecture. Trunks, MPO counts, pathways, switch positions, and PDU circuit landings decide whether the next cell is additive or a recable. Pre-build and reserve where the first transition is already visible. Do not pull a Day-1 plant that can only ever be Day-1. - Trunks: Structured trunks with spare fibers, not a jumper for every live port. A spare in the trunk is cheaper than a night in the hall. - MPO counts: Count lanes for the optics class you will grow into, not only the optic seated today. Mixing MPO-12 leftovers into an MPO-16 plant is how the next leaf stays dark. - Pathways: Reserved basket and tray for the next trunk. If the only path is full, additive growth is already a transition. - Switch positions: U and power for the next leaf, with the rail map already drawn. A free ToR slot with no trunk landing is not reserve. - PDUs: Circuit count and landing for the next node or leaf. Cabling the packet path while the power path is exhausted is how a hall dead-ends. - Pre-build / reserve: Justified when Day-1 already names the first transition. Not justified as dark fiber for a fantasy campus. Write the gate. See /expansion. Canonical: https://neuronplant.com/cabling#tomorrow ## AI platform Serving, RAG services, NIM / runtime integration, and the data path from sources through ingest, train or RAG, Docker/Kubernetes, KV / prefix cache, operate / observe, handover. HTML: https://neuronplant.com/applications Split file: https://neuronplant.com/llms/pipeline.txt # NeuronPlant applications / platform Accelerators without a platform are inventory. The factory includes serving, RAG services, NIM / runtime integration, and the path from data to workload. NeuronPlant delivers AI factories. After power, cooling, racks and accelerators exist, there is still a platform: serving, RAG services, NIM / runtime integration, plus a data path from sources through ingest, storage, train/fine-tune or RAG, accelerator runtime (Docker images, Kubernetes cell), inference serving (NVIDIA NIM from public docs), KV / prefix cache on large serving, applications, then operate / observe. NeuronPlant owns that platform path as part of the factory delivery: architecture, integration onto the plant (network planes, storage tiers, NIM/profile sizing from NVIDIA public documentation, KV / prefix cache as a serving-plane constraint), and handover. NeuronPlant does not claim to write every model, is not a generic software ISV, and is not an NVIDIA Cloud Partner. ## Stages ### 01 Data sources Where bytes originate: operational systems, files, streams, instruments, model registries. Plant: Lands on the North-South / front-end plane, not on the East-West compute rails. ### 02 Ingest Controlled intake: validation, classification, residency, and a path that can be audited. Plant: Sized to nightly rebuilds or continuous feed. Must not contend with checkpoint burst. ### 03 Storage Tiers, not a single NAS. Documents, datasets, checkpoints, indexes and weights have different machines. Plant: Storage/data fabric off the East-West compute rails when dedicated. Object, parallel FS, and node-local NVMe are different jobs. Do not auto-label this North-South or East-West without that context. ### 04A Train / fine-tune Distributed training or adaptation on a coupled accelerator domain. Checkpoints are a storage event. Plant: East-West compute fabric (Quantum-X or Spectrum-X) plus HPS sequential GB/s. See SuperPOD public RAs. ### 04B RAG / retrieval Document store plus a vector index. Ingest and query are IOPS and metadata, not training scratch. Plant: Keep the index off the training filesystem. North-South / front-end Ethernet is usually enough if the index is close. That is not the East-West compute fabric. ### 05 Accelerator runtime The cluster that actually runs the job. Workloads ship as Docker images. Kubernetes schedules, restarts, and places them. Placement, isolation, and a scheduler that knows the fabric. Plant: A starting cell is three masters and three workers. Masters may be VMs. Workers that hold accelerators are usually bare metal. Observe sits on this plane. Same accelerators as the hall. ### 06 Inference platform Serving plane. NVIDIA NIM is the cited runtime for profiled LLM serving. Precision, TP size and LoRA come from NVIDIA's public support matrix. On a large job, KV and prefix cache are a memory hierarchy, not the weight disk. Plant: Local NVMe for images and weights. KV lives in HBM first, then host overflow, then a shared prefix cache across the serving plane. Size that hierarchy separately from the profile disk. ### 07 Applications Where the factory meets the enterprise: assistants, tools, batch, APIs. Ownership of the app stays with the operator. Plant: North-South front-end path, identity, and a handover that names interfaces. Not a NeuronPlant SaaS. Not management/OOB. ## KV cache / large serving KV cache is a serving constraint, not a storage array. On transformer inference, each new token attends to the keys and values of every prior token. Those K/V tensors are cached so the runtime does not recompute the whole prefix on every decode step. A small job keeps KV in accelerator HBM. GPU is the usual home. That is enough. A large workload is different. Long context, many concurrent sessions, and shared system prompts or RAG prefixes fill HBM. The plant then needs a memory hierarchy: HBM → RAM → NVMe → remote, plus a prefix cache and KV-aware routing so the fleet does not recompute the same context. Prefill writes the cache. Decode reads it. That split changes how you size HBM, host memory, the inference fabric, and sometimes a fast tier next to the accelerators. This is not a storage array. It is not a training checkpoint pool. The constraint is the job: serving plane, prefix cache, prefill versus decode, HBM versus host. The question at scale: do not recompute context across a growing fleet. Hierarchy is HBM → RAM → NVMe → remote. Too heavy for the homepage. It lives here. - On-device HBM: First home for active KV. Fine until context length and concurrency fill the device. - Host overflow / RAM: When HBM fills, KV spills to host memory. That is a memory and fabric size, not a new LUN on the training filesystem. - NVMe: Next overflow when host RAM is not enough. Still a serving-plane tier, still not the training filesystem. - Remote: A shared cache off the node when the fleet should not keep rebuilding the same prefix. Latency and fabric become the spec. - Local vs prefix: Per-request KV stays local. Shared system prompts and RAG prefixes belong in a prefix cache so the fleet does not recompute the same context. - KV-aware routing: Send the next turn to the replica that already holds the cache. Round-robin across a growing fleet is how you pay prefill again. ## RAG is a data architecture, not a vector database. Retrieval is a path with named data states. Source → landing → parse → process → chunk → embed → index → retrieve → LLM. A vector engine is one station. It is not the plant. - 01 Source (Authoritative corpus. Systems of record, files, streams.) Name where truth lives. Size storage from this TB, not from the index. - 02 Landing (As-received bytes. Copy or reference, written down.) Controlled intake onto the plant. Must not contend with checkpoint burst. - 03 Parse (Structure extracted. Layout, tables, attachments.) CPU workers, not accelerator nodes. Failures here poison every downstream state. - 04 Process (Cleaned, classified, residency applied.) ACL and retention start here, not at the chatbot. - 05 Chunk (Retrieval units. Overlap and size are a design.) Changing chunking is a re-embed. Treat it as a plant event. - 06 Embed (Vectors for those chunks, versioned with the model.) An embedding-model change rebuilds this state and the index. GPU is the common accelerator. - 07 Index (Queryable ANN / keyword / hybrid.) IOPS and metadata. Keep off the training filesystem. - 08 Retrieve (Candidates for this query, under ACL.) Latency to the serving plane. Delete and ACL must propagate here, not only in the source. - 09 LLM (Generated answer. Not a source of truth.) Inference plus KV. The corpus does not live in the prompt. ## If you cannot rebuild it, you do not own it. For every state: where it lives, how long, what is authoritative, how it rebuilds, how it is backed up, how it is deleted, and who is allowed to delete it. - Where: Which tier and which plane. Landing is not the index. The index is not the LLM. - How long: Retention per state. Hot retrieve, legal hold, and disposable rebuilds are different clocks. - Authoritative: Source of record. Embeddings are derived. Chat transcripts are not the corpus. - Rebuild: From which upstream state, in what window, with what compute. Write it before the first index. - Backup: What is cheaper to protect than to rebuild. Derived vectors are often the latter. - Delete / who deletes: A named owner. Source delete is not index delete until it has propagated. - Embedding-model change: A plant event: re-chunk or re-embed, new index, dual-run window. Not a config toggle on the chatbot. - ACL / delete propagation: Permissions and erasures have to reach retrieve, not only the system of record. Otherwise the factory leaks. ## Runtime / Operate Canonical page: /runtime Docker: the workload ships as an image. Same artifact from lab to factory. Kubernetes: schedule, restart, and place those containers. A fleet that is ssh-and-excel will not stay up. Starting cell: three masters (VM-capable) and three workers (workers that hold accelerators are usually bare metal). Fluent Bit is the example log shipper. ## Fork Train / fine-tune and RAG are different I/O machines that can share a hall. Mixing them on one filesystem is how jobs stall. ## In scope - Architecture of the data-to-workload path - Integration onto the plant: network planes, storage tiers, rack placement - NVIDIA NIM / profile sizing from NVIDIA public documentation - KV / prefix cache as a serving-plane constraint on large inference, including HBM → RAM → NVMe → remote and KV-aware routing - RAG as a data path with named states, sized from source data, with a rebuildable lifecycle - Docker images and a Kubernetes cell (3 masters, 3 workers) as the factory runtime - Observe / monitor as part of that plane, with Fluent Bit as the example shipper - Orchestration and MLOps as interfaces, not as a NeuronPlant product - Handover: runbooks, boundaries, and who operates which plane ## Out of scope - Writing every model the customer will run - A generic software ISV or SaaS application suite - NVIDIA partnership, Cloud Partner, or NVOnline status - Named customer logos or invented case studies ## Related factory templates - Private AI (/ai-factories/private-ai): Enterprise sources, controlled ingest, RAG, NIM serving, internal applications. - Training cluster (/ai-factories/training): Datasets, staging, coupled fabric, checkpoints, scheduler, experiment platform. - AI-as-a-Service (/ai-factories/ai-as-a-service): Onboarding, tenancy, accelerator pools, shared fabric, control plane, customer workloads. NVIDIA NIM support matrix: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/applications # KV cache is a serving constraint, not a storage array. On transformer inference, each new token attends to the keys and values of every prior token. Those K/V tensors are cached so the runtime does not recompute the whole prefix on every decode step. A small job keeps KV in accelerator HBM. GPU is the usual home. That is enough. A large workload is different. Long context, many concurrent sessions, and shared system prompts or RAG prefixes fill HBM. The plant then needs a memory hierarchy: HBM → RAM → NVMe → remote, plus a prefix cache and KV-aware routing so the fleet does not recompute the same context. Prefill writes the cache. Decode reads it. That split changes how you size HBM, host memory, the inference fabric, and sometimes a fast tier next to the accelerators. This is not a storage array. It is not a training checkpoint pool. The constraint is the job: serving plane, prefix cache, prefill versus decode, HBM versus host. The question at scale: do not recompute context across a growing fleet. Hierarchy is HBM → RAM → NVMe → remote. Too heavy for the homepage. It lives here. - On-device HBM: First home for active KV. Fine until context length and concurrency fill the device. - Host overflow / RAM: When HBM fills, KV spills to host memory. That is a memory and fabric size, not a new LUN on the training filesystem. - NVMe: Next overflow when host RAM is not enough. Still a serving-plane tier, still not the training filesystem. - Remote: A shared cache off the node when the fleet should not keep rebuilding the same prefix. Latency and fabric become the spec. - Local vs prefix: Per-request KV stays local. Shared system prompts and RAG prefixes belong in a prefix cache so the fleet does not recompute the same context. - KV-aware routing: Send the next turn to the replica that already holds the cache. Round-robin across a growing fleet is how you pay prefill again. Canonical: https://neuronplant.com/applications#kv-cache # RAG is a data architecture, not a vector database. Retrieval is a path with named data states. Source → landing → parse → process → chunk → embed → index → retrieve → LLM. A vector engine is one station. It is not the plant. # If you cannot rebuild it, you do not own it. For every state: where it lives, how long, what is authoritative, how it rebuilds, how it is backed up, how it is deleted, and who is allowed to delete it. - Where: Which tier and which plane. Landing is not the index. The index is not the LLM. - How long: Retention per state. Hot retrieve, legal hold, and disposable rebuilds are different clocks. - Authoritative: Source of record. Embeddings are derived. Chat transcripts are not the corpus. - Rebuild: From which upstream state, in what window, with what compute. Write it before the first index. - Backup: What is cheaper to protect than to rebuild. Derived vectors are often the latter. - Delete / who deletes: A named owner. Source delete is not index delete until it has propagated. - Embedding-model change: A plant event: re-chunk or re-embed, new index, dual-run window. Not a config toggle on the chatbot. - ACL / delete propagation: Permissions and erasures have to reach retrieve, not only the system of record. Otherwise the factory leaks. Canonical: https://neuronplant.com/applications#rag # NeuronPlant runtime / operate The factory runs as a fleet of images, not as a snowflake hall. After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane. ## Why Docker The workload ships as an image. Serving, ingest, RAG, NIM runtime, plant services. The hall is not a unique snowflake OS. Same artifact from lab to factory. ## Why Kubernetes Schedule, restart, and place those containers across the plant. Accelerator nodes, CPU nodes, and plant services in one control plane. Not because cloud native is a slogan. An AI factory is a fleet. A fleet that is ssh-and-excel will not stay up. ## What Kubernetes requires A starting cell: three masters and three workers. That is a plant constraint, not a certification syllabus. - Three masters: Three masters so etcd and the API stay up if one dies. That is why 3, not 1. - Three workers: Three workers as a starting factory cell. The cell can grow. Accelerator workers are not the same machine class as the masters. - Masters may be VMs: Masters may be virtual machines. Workers that hold accelerators are usually bare metal. The plant does not require three physical master servers. - Plant jobs, not a SKU list: Network, storage classes, and an accelerator device plugin are jobs on the plane. They are not a product catalog we refresh every week. ## Observability Logs, metrics, and traces are part of the plane. Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards. ## Kubernetes as production infrastructure Not K8s plus vLLM as a slogan. A factory control plane with CPU workers, accelerator workers, and the platform services that must not sit on the expensive nodes. - Control plane: API, etcd, schedulers. Masters may be VMs. This plane stays up if a GPU node dies. - CPU workers: Ingest, parse, RAG services, registries, identity, plant jobs. Sized as a cell, grown as a gate. See /expansion. - Accelerator workers: GPU is the common case. TPU and other ASICs when that is the job. Bare metal. Device plugin. Not a place to park Prometheus. - Registry: The images the factory actually runs. On the management plane, not on an HBM node. - Secrets: A named store and a rotation path. Not files on the GPU worker. - Identity: Who may schedule, who may retrieve, who may delete. RAG ACL starts here and has to reach the index. - RAG services: Parse, embed jobs, retrievers. CPU first. Accelerators only for the embed/infer step that needs them. - Keep platform off accelerator nodes: Do not park registry, logging, identity, or the control plane on expensive accelerator workers. That is how a fleet starves itself. ## In scope - Docker as the image that moves serving, ingest, RAG, NIM, and plant services from lab to factory - Kubernetes as the control plane that schedules that fleet across accelerator, CPU, and plant nodes - A starting cell of three masters and three workers, with masters allowed to be VMs - Observability as part of the plane, with Fluent Bit as the example log shipper - Production Kubernetes: control plane, CPU workers vs accelerator workers, registry, secrets, identity, RAG services - Platform services kept off expensive accelerator nodes ## Out of scope - A CNCF certification syllabus or weekly operator-version pin list - A CNI, service-mesh, or dashboard catalog we would have to refresh every week - Requiring three physical master servers when VMs will hold quorum - Treating ssh-and-excel as the runbook for an accelerator hall - K8s plus vLLM as the whole platform story - Named TTFT / ITL numbers as a NeuronPlant product Canonical page: /runtime. Also on /what-we-build#operate and /applications#runtime. Canonical: https://neuronplant.com/runtime # Kubernetes as production infrastructure Not K8s plus vLLM as a slogan. A factory control plane with CPU workers, accelerator workers, and the platform services that must not sit on the expensive nodes. - Control plane: API, etcd, schedulers. Masters may be VMs. This plane stays up if a GPU node dies. - CPU workers: Ingest, parse, RAG services, registries, identity, plant jobs. Sized as a cell, grown as a gate. See /expansion. - Accelerator workers: GPU is the common case. TPU and other ASICs when that is the job. Bare metal. Device plugin. Not a place to park Prometheus. - Registry: The images the factory actually runs. On the management plane, not on an HBM node. - Secrets: A named store and a rotation path. Not files on the GPU worker. - Identity: Who may schedule, who may retrieve, who may delete. RAG ACL starts here and has to reach the index. - RAG services: Parse, embed jobs, retrievers. CPU first. Accelerators only for the embed/infer step that needs them. - Keep platform off accelerator nodes: Do not park registry, logging, identity, or the control plane on expensive accelerator workers. That is how a fleet starves itself. Canonical: https://neuronplant.com/runtime#plane ## Delivery records Anonymized factory records. No invented customer names. HTML: https://neuronplant.com/projects Split file: https://neuronplant.com/llms/project.txt # Private AI Factory / Financial Sector Client record: Financial Institution / Israel Deploy a private AI environment while maintaining data sovereignty. ## Scale - 32 GPUs - 400/800G fabric - High-performance storage - Liquid-cooled infrastructure ## Scope - Architecture - Power - Cooling - Compute - Network - Cabling - Storage - Deployment Outcome: Operational private AI environment with defined expansion capacity. Canonical: https://neuronplant.com/projects/private-ai-financial # AI Training Cluster / Research Client record: Research Organization Stand up a shared training cluster that researchers can schedule without turning the facility into an experiment. ## Scale - 64 GPUs - InfiniBand fabric - Parallel filesystem - Direct liquid cooling ## Scope - Feasibility - Architecture - Cooling - Network - Cabling - Storage - Validation Outcome: A validated training domain with documented power, cooling, and job-level performance envelopes. Canonical: https://neuronplant.com/projects/training-cluster-research # GPU Service Cell / Service Provider Client record: Global Technology Company Design a repeatable GPU cell that can be dropped into existing colo halls and sold as capacity. ## Scale - 128+ GPUs / cell - 800G Ethernet - Multi-tenant storage - Liquid-ready rows ## Scope - Architecture - Site - Power - Cooling - Network - Cabling - Platform integration Outcome: A cell specification with BOM, rack elevations, and a second-hall expansion path. Canonical: https://neuronplant.com/projects/gpu-cloud-cell # High-density GPU hall / structured cabling failure Client record: Operator withheld The operator bought a large CPU and GPU fleet. Structured cabling was let to another firm as a fit-out. The hall has been offline for a long time. NeuronPlant was not the cabling contractor. ## Scale - High-density GPU hall - Large CPU + GPU buy - 400/800G class fabric - Third-party structured cabling ## Scope - Not the original cabling contractor - Plant audit if engaged - Polarity / IL / OTDR map - As-built reconstruction - Per-fabric recable sequence - Test gates before optics reseat Outcome: Still a cautionary record, not a NeuronPlant win. Compute sat while the original plant remained the plant of record. Recovery is a specified recable with test records. It is not a firmware patch. This is a cautionary record, not a NeuronPlant win. No customer name. Canonical: https://neuronplant.com/projects/gpu-hall-passive-plant ## Customer specifications Held behind the same written-approval gate as named customers. HTML: https://neuronplant.com/customers Split file: https://neuronplant.com/llms/customer-spec.txt _No published records._