01 / Unmanaged distribution
Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks.
Nine layers is the canonical plant architecture: the factory NeuronPlant delivers as one project. Structured cabling and passive infrastructure is not a footnote under networking. The AI platform is not a software product we write. Operate is how the factory stays visible after handover. Site is the hall we occupy: a partner room, not a campus we pour.
NP-SYS-09 / Interconnected plant
Accelerator plant. The hall is a supplier
The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it.
From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job.
Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies.
Accelerators as an engineered cluster. GPU is the common case, not the only one.
East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South.
Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark.
Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure.
The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model.
The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog.
NP-MAP / AI Factory System View
The numbered plant is nine layers. This drawing is a dependency stack: grid, hall, plant, workload, and applications. It is not another sequence of an AI factory.
Dependency map. Facility constrains compute. Cabling lands the fabrics.
NP-CFG / Starting plant
Conceptual configuration of the GPU / TPU / ASIC cell. Site is the hall we occupy, not a campus we construct. NeuronPlant takes this from architecture through purchase, deployment, validation, and handover.
This chooser sizes a GPU plant. GPU is the common case. TPU and other accelerators / ASICs are the same compute layer. We do not list a chip catalog. Site is the hall we occupy, not a campus we construct.
Accelerator plant · conceptual
01 / Site
The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it.
02 / Power
From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job.
A PDU is how the rack gets power. Managed means the PDU has a management plane: metering (rack, phase, or outlet), network, often remote outlet control.
Managed is not automatic. Specify the job: unmanaged distribution, metered, or switched at the outlet. This is not a vendor catalog. Managed PDU and unmanaged PDU are the durable words.
NP-PWR-PDU / Dual-cord rack
Unmanaged · metered · switched. Specify the job.
01 / Unmanaged distribution
Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks.
02 / Metered
See kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. Justified when the board or BMS does not already meter at this grain, or when the hall is multi-tenant or billed.
03 / Switched / outlet-level
Remote outlet control. On dual-cord GPU and compute racks this is a foot-gun: you do not power-cycle a training node from a PDU web UI. Specify it only when that job is real.
03 / Cooling
Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies.
04 / Compute
Accelerators as an engineered cluster. GPU is the common case, not the only one.
The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. NVIDIA DGX and GB-class remain how we size a GPU plant. Specs are NVIDIA's. The room we occupy, the loop, and the fabric are the design. We do not maintain a weekly chip catalog.
05 / Network
East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South.
06 / Structured cabling & passive infrastructure
Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark.
NP-CBL / Fabric planes
East-West · North-South · Management. Cabling builds the plants.
DWG-CBL-01 · laying out
EW / East-West
Hall-internal scale-out. Compute fabric first. Storage/data fabric on its own plane when dedicated. Not user traffic.
01 / East-West Compute Fabric
MPO/MTP trunks, 400/800G optics, rail-aligned
GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.
02 / Storage / Data Fabric
Separate MPO or LC plant onto the storage leaf when dedicated
Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails.
NS / North-South
Users, applications, API gateway/load balancer, enterprise/core, firewalls. Application-facing Ethernet. Not another GPU rail.
03 / North-South / Front-End Fabric
LC or lower-count MPO, 100G-class typical
Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O.
MGMT / Management
BMC, management, operational control. Never classify as application North-South.
04 / Management / OOB Fabric
Cat6A or dedicated 1/10G optics. Own patch field.
BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark.
NP-CBL-MED / Media and reach
DAC · AEC · AOC · MMF · SMF
DAC
In-rack / adjacent. Typically sub-3 m at 400/800G.
ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend.
Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly.
AEC
Active copper. Longer than DAC, still copper.
In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch.
Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail.
AOC
In-row to in-hall. Fiber with the optics already on the ends.
When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut.
Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant.
MMF (OM4 / OM5)
In-hall SR-class. MPO/MTP trunks.
Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site.
Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested.
SMF (OS2)
Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G.
Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU.
Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone.
NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF.
Full plant, media, practice and failure modes on Structured Cabling & Passive Infrastructure.
07 / Storage
Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure.
08 / Platform
The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model.
NP-PIPE / Data to workload
Schematic · not a product flowchart
Large serving includes a KV / prefix cache hierarchy. That is a memory job on the serving plane, not another storage array. Definition on Applications / Platform.
09 / Operate
The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog.
After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane. Logs, metrics, and traces are part of the plane. Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards. Production Kubernetes keeps CPU workers, registry, secrets, identity, and RAG services off the expensive accelerator nodes. Runtime / production plane.
The engagement record itself is a client program, not a spreadsheet. Preview on Client portal. Not live.
This is not the commercial engagement. The customer journey is Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover. These steps are the engineering work inside that delivery: feasibility of the room, detailed design, BOM, install, commission, and observe.
Plant engineering sequence. Distinct from the commercial journey on the homepage: Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover.
01 · Discover
02 · Discover
03 · Design
04 · Detail
05 · Build
06 · Build
07 · Commission
08 · Observe
How we bid sits with the firm, not in the BOM. Integrity before a steered RFP: About / Integrity. Additive versus transition, expansion gates, and the two Day-1 money paths: Expansion / Day-1.
NP / Intake
One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.