Skip to main content
NP-WWB/What we build

The GPU, TPU, and ASIC plant.

Nine layers is the canonical plant architecture: the factory NeuronPlant delivers as one project. Structured cabling and passive infrastructure is not a footnote under networking. The AI platform is not a software product we write. Operate is how the factory stays visible after handover. Site is the hall we occupy: a partner room, not a campus we pour.

NP-SYS-09 / Interconnected plant

01

Site

The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it.

  • · Hall / colo cage we place into
  • · Partner room or custom-system supplier
  • · Rack planning in that room
  • · Floor loading
02

Power

From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job.

  • · Hall-presented power
  • · Capacity planning for the cell
  • · UPS as the room allows
  • · PDU / busway in the cell
03

Cooling

Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies.

  • · Liquid cooling
  • · CDU design
  • · Facility water requirements
  • · Heat rejection
04

Compute

Accelerators as an engineered cluster. GPU is the common case, not the only one.

  • · GPU (common case)
  • · TPU
  • · Other accelerators / ASICs
  • · NVIDIA DGX / GB-class when the plant is GPU
05

Network

East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South.

  • · NVIDIA Quantum-X / Spectrum-X East-West compute
  • · Storage / data fabric (dedicated when required)
  • · North-South / front-end fabric
  • · Management / OOB fabric
06

Structured cabling & passive infrastructure

Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark.

  • · Fiber and MPO/MTP trunks (MMF / SMF, 400/800G optics)
  • · Patch panels, cross-connects, rack-to-rack
  • · Cable pathways and cable management
  • · Copper and passive optical infrastructure
07

Storage

Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure.

  • · High-performance storage
  • · Throughput / IOPS / metadata
  • · AI datasets
  • · Training storage
08

Platform

The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model.

  • · Platform architecture
  • · Ingest and dataset path
  • · Train / fine-tune vs RAG
  • · Accelerator runtime placement
09

Operate

The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog.

  • · Docker images
  • · Kubernetes cell (3 masters + 3 workers)
  • · Masters may be VMs
  • · CPU workers vs accelerator workers

NP-MAP / AI Factory System View

How the factory meets utility, workload, and applications.

The numbered plant is nine layers. This drawing is a dependency stack: grid, hall, plant, workload, and applications. It is not another sequence of an AI factory.

  • Grid
  • Power
  • Cooling
  • AI Data Center
  • Accelerated ComputeGPU · TPU · ASIC
  • Structured Cabling
  • AI Fabrics
  • Storage & Data
  • AI Platform
  • AI Workloads
  • Applications & APIs

NP-CFG / Starting plant

A starting cell for a NeuronPlant delivery

Conceptual configuration of the GPU / TPU / ASIC cell. Site is the hall we occupy, not a campus we construct. NeuronPlant takes this from architecture through purchase, deployment, validation, and handover.

Workload
Scale
GPU plant (NVIDIA)

This chooser sizes a GPU plant. GPU is the common case. TPU and other accelerators / ASICs are the same compute layer. We do not list a chip catalog. Site is the hall we occupy, not a campus we construct.

Accelerator plant · conceptual

Site
2 racks in a partner hall path.
Power
120 kW. Dual-cord: unmanaged or metered-not-switched. Do not outlet-cycle a training node.
Cooling
Air or rear-door
Compute
8 × NVIDIA DGX B300 (64 GPUs)
Network
Quantum InfiniBand East-West compute fabric (X800 or Quantum-2 class)
Structured cabling
Rail-aligned MPO, Type-B polarity map, EW compute / storage / OOB on separate plants
Storage
HPS >40 GB/s per node (NVIDIA SuperPOD B300 RA) + checkpoint tier
Platform
Datasets → staging → train fabric → checkpoints → scheduler
Operate
3+3 cell. Observe the fabric and the job. Fluent Bit at the node edge.
Start from this concept →

01 / Site

Site

The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it.

  • Hall / colo cage we place into
  • Partner room or custom-system supplier
  • Rack planning in that room
  • Floor loading
  • U / weight / clearance
  • kW and heat the hall can present
  • Path in / meet-me

02 / Power

Power

From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job.

  • Hall-presented power
  • Capacity planning for the cell
  • UPS as the room allows
  • PDU / busway in the cell
  • Managed vs unmanaged PDU
  • Rack power distribution
  • Redundancy architecture

A PDU is how the rack gets power. Managed means the PDU has a management plane: metering (rack, phase, or outlet), network, often remote outlet control.

Managed is not automatic. Specify the job: unmanaged distribution, metered, or switched at the outlet. This is not a vendor catalog. Managed PDU and unmanaged PDU are the durable words.

NP-PWR-PDU / Dual-cord rack

Unmanaged · metered · switched. Specify the job.

DUAL-CORD GPU RACK / A-B FEEDSBUSWAY ABUSWAY BMGMT PLANEPDU Aunmanaged unless specifiedPDU Bdual-cord, not a web UIGPU NODEDual-cord. Do not outlet-cyclea training node from the PDU.only if the job is meteringor switching. Not by default.

01 / Unmanaged distribution

Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks.

02 / Metered

See kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. Justified when the board or BMS does not already meter at this grain, or when the hall is multi-tenant or billed.

03 / Switched / outlet-level

Remote outlet control. On dual-cord GPU and compute racks this is a foot-gun: you do not power-cycle a training node from a PDU web UI. Specify it only when that job is real.

03 / Cooling

Cooling

Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies.

  • Liquid cooling
  • CDU design
  • Facility water requirements
  • Heat rejection
  • Chillers
  • Cooling redundancy

04 / Compute

Compute

Accelerators as an engineered cluster. GPU is the common case, not the only one.

  • GPU (common case)
  • TPU
  • Other accelerators / ASICs
  • NVIDIA DGX / GB-class when the plant is GPU
  • Cluster architecture
  • Capacity planning

The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. NVIDIA DGX and GB-class remain how we size a GPU plant. Specs are NVIDIA's. The room we occupy, the loop, and the fabric are the design. We do not maintain a weekly chip catalog.

05 / Network

Network

East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South.

  • NVIDIA Quantum-X / Spectrum-X East-West compute
  • Storage / data fabric (dedicated when required)
  • North-South / front-end fabric
  • Management / OOB fabric
  • Converged vs dedicated compute and storage
  • 800G fabrics

06 / Structured cabling & passive infrastructure

Structured cabling & passive infrastructure

Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark.

  • Fiber and MPO/MTP trunks (MMF / SMF, 400/800G optics)
  • Patch panels, cross-connects, rack-to-rack
  • Cable pathways and cable management
  • Copper and passive optical infrastructure
  • Separate trays for EW compute, storage/data, N-S, and OOB
  • Physical connectivity reserved for expansion
  • Labeling, IL / OTDR, polarity test, as-builts

NP-CBL / Fabric planes

East-West · North-South · Management. Cabling builds the plants.

DWG-CBL-01 / NOT FOUR EQUAL JOBSSTRUCTURED CABLING BUILDS THESE PLANTSENTERPRISE / USERS / APPLICATIONSAPI GATEWAY / LB · FIREWALLS · CORENORTH-SOUTH / FRONT-END FABRICFRONT-END / APP / ENTERPRISEAI PLATFORM / DGX ENVIRONMENTNOT THE HALL · THE ACCELERATOR CELLEAST-WEST COMPUTE · QUANTUM-X / SPECTRUM-XGPU-TO-GPUDGX-TO-DGX · RDMASTORAGE / DATA · DEDICATED · COMPUTE ↔ HPSMGMT / OOB · NOT N-SDGXGPU rail 1DGXGPU rail 2DGXGPU rail 3DGXGPU rail 4OVERHEAD BASKETS / HOW THE PLANTS ARE BUILT / NOT A FIFTH FABRICEW · TRAY A · COMPUTE MPO · RAIL-ALIGNEDEW · TRAY B · STORAGE / DATA · OFF THE COMPUTE RAILSNS · TRAY C · FRONT-ENDMGMT · TRAY D · OOB / BMCN-S IS USERS AND APPSE-W IS GPU / DGX SCALE-OUTUNDERFLOOR = POWER / LIQUID · OVERHEAD = FIBER · SLACK AT RACK AND FRAMESTORAGE AND OOB INDEPENDENT

DWG-CBL-01 · laying out

EW / East-West

Hall-internal scale-out. Compute fabric first. Storage/data fabric on its own plane when dedicated. Not user traffic.

01 / East-West Compute Fabric

MPO/MTP trunks, 400/800G optics, rail-aligned

GPU-to-GPU, DGX-to-DGX, scale-out. RDMA/RoCE or InfiniBand. NVIDIA Spectrum-X or Quantum-X as the high-speed EW compute fabric. NCCL, all-reduce. One polarity map. One tray.

02 / Storage / Data Fabric

Separate MPO or LC plant onto the storage leaf when dedicated

Compute to high-performance storage: datasets and checkpoints. Own internal data plane when physically separate. Do not fold it into a generic East-West compute line. Do not auto-label it North-South or East-West without that context. NVIDIA SuperPOD keeps this off the compute rails.

NS / North-South

Users, applications, API gateway/load balancer, enterprise/core, firewalls. Application-facing Ethernet. Not another GPU rail.

03 / North-South / Front-End Fabric

LC or lower-count MPO, 100G-class typical

Users, apps, API gateway/LB, enterprise/core, application-facing Ethernet. Not just another rail next to the GPU rails. Must not share a tray or a cassette with training I/O.

MGMT / Management

BMC, management, operational control. Never classify as application North-South.

04 / Management / OOB Fabric

Cat6A or dedicated 1/10G optics. Own patch field.

BMC, management, operational control. Lights-out, firmware, serial, leak sensors. Never application North-South. If this plant dies inside a compute bundle, you cannot even see why the hall is dark.

NP-CBL-MED / Media and reach

DAC · AEC · AOC · MMF · SMF

DAC<3 mAECin-row CuAOCin-row fiberMMF MPOin-hall SRSMF MPOCIN 500 mREACH CLASS · NOT A CATALOG400/800G optics follow the plant. Polarity is Type A / Type B as a drawing.

DAC

In-rack / adjacent. Typically sub-3 m at 400/800G.

ToR to NIC when the geometry is short and the rear of the rack can take the copper bulk and the bend.

Avoid: Row-to-row. Stuffed chimneys on air-cooled HGX. Any run you will service weekly.

AEC

Active copper. Longer than DAC, still copper.

In-row when you need retimed copper and cannot land fiber yet. Treat power and heat as part of the cable, not as a free lunch.

Avoid: Campus. Underfloor spaghetti. Mixing AEC with untested MPO adapters on the same rail.

AOC

In-row to in-hall. Fiber with the optics already on the ends.

When you want optical reach without field-terminated MPO. Good for known lengths. Bad if you need to recut.

Avoid: A plant you will grow by splicing. AOC is a SKU, not a structured plant.

MMF (OM4 / OM5)

In-hall SR-class. MPO/MTP trunks.

Leaf to leaf inside a hall when the SR optic and the MPO count match. Polarity is a design, not a guess on site.

Avoid: CIN / campus distances. Mixing OM3 leftovers into an OM4 trunk and calling it tested.

SMF (OS2)

Hall to hall, and CIN-class. DR/FR/LR optics at 400/800G.

Spine, meet-me, storage across halls. NVIDIA's public DSX overview puts a 500 m optical reach limit on the cluster interconnect. That is a fiber plant number, not a switch SKU.

Avoid: Unspecified polarity. Unlabeled MPO-16 vs MPO-12. No OTDR on the backbone.

NVIDIA's public DSX facilities overview states a 500 m optical reach limit on the cluster interconnect (CIN) spine. Hall placement, meet-me, and SMF plant all have to land inside that class. A switch that can do 800G does not help if the fiber path is 700 m of unspecified MMF.

Full plant, media, practice and failure modes on Structured Cabling & Passive Infrastructure.

07 / Storage

Storage

Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure.

  • High-performance storage
  • Throughput / IOPS / metadata
  • AI datasets
  • Training storage
  • RAG sized from source data
  • Checkpointing
  • Backup and data lifecycle

08 / Platform

Platform

The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model.

  • Platform architecture
  • Ingest and dataset path
  • Train / fine-tune vs RAG
  • Accelerator runtime placement
  • NVIDIA NIM profile sizing
  • KV / prefix cache (large serving)
  • Orchestration / MLOps interfaces
  • Enterprise application handover

NP-PIPE / Data to workload

Schematic · not a product flowchart

PLANT RAILDWG-PIPE-01GRIDPOWERCOOLINGHALLCOMPUTECABLE / STOINTEGRATION PLANE · NETWORK PLANES · STORAGE TIERS · NIM PROFILE SIZING01DATA SOURCESOps · files · streams02INGESTValidate · classify03STORAGEObject · HPS · NVMe04ATRAIN / FINE-TUNECoupled compute domain04BRAG / RETRIEVALDocs + vector index05RUNTIMEDocker · Kubernetes · observe06INFERENCE / NIMProfile · TP · LoRA07APPLICATIONSEnterprise edgeAccelerators without this path are inventory. The factory includes the path from data to workload.NeuronPlant: architecture, plant integration, NIM sizing from NVIDIA public docs, handover. Not a model shop.FORK · TRAIN / FINE-TUNE OR RAG

Large serving includes a KV / prefix cache hierarchy. That is a memory job on the serving plane, not another storage array. Definition on Applications / Platform.

09 / Operate

Operate

The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog.

  • Docker images
  • Kubernetes cell (3 masters + 3 workers)
  • Masters may be VMs
  • CPU workers vs accelerator workers
  • Observe / monitor
  • Fluent Bit (example shipper)
  • Runbooks after handover

After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane. Logs, metrics, and traces are part of the plane. Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards. Production Kubernetes keeps CPU workers, registry, secrets, identity, and RAG services off the expensive accelerator nodes. Runtime / production plane.

The engagement record itself is a client program, not a spreadsheet. Preview on Client portal. Not live.

NP-04/Inside the delivery

How the plant is engineered

This is not the commercial engagement. The customer journey is Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover. These steps are the engineering work inside that delivery: feasibility of the room, detailed design, BOM, install, commission, and observe.

Plant engineering sequence. Distinct from the commercial journey on the homepage: Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover.

Discover
Design
Detail
Install
Commission
Observe

01 · Discover

Discovery

  • Business objectives
  • AI workloads
  • Models
  • Dataset size
  • Performance requirements
  • Growth expectations

02 · Discover

Feasibility

  • Power availability
  • Cooling availability
  • Facility limitations
  • Connectivity
  • Space
  • Budget

03 · Design

Architecture

  • Hall we occupy
  • Power architecture
  • Cooling architecture
  • Compute architecture
  • Network fabric
  • Structured cabling & passive infrastructure
  • Storage architecture
  • AI platform
  • Runtime / observe

04 · Detail

Engineering

  • Detailed design
  • Rack layouts
  • Electrical requirements
  • Cooling requirements
  • Cabling / polarity map
  • Platform integration
  • Bill of materials

05 · Build

Procurement & Coordination

  • Purchase against the architecture
  • Vendors
  • Hall partner / custom-system supplier
  • OEMs
  • Coordination
  • Logistics

06 · Build

Deployment

  • Installation
  • Integration
  • Configuration
  • Testing

07 · Commission

Validation

  • Performance testing
  • Network validation
  • Structured cabling test (IL, polarity, OTDR)
  • Storage validation
  • Cooling validation
  • Failover testing

08 · Observe

Handover

  • Documentation
  • Training
  • Operational procedures
  • Observe / monitor
  • Runtime cell (Docker / K8s)
  • Support model

How we bid sits with the firm, not in the BOM. Integrity before a steered RFP: About / Integrity. Additive versus transition, expansion gates, and the two Day-1 money paths: Expansion / Day-1.

NP / Intake

Tell us what you need to achieve with AI.

One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.

Start a Project →