# NeuronPlant / Plant layers Site, power, cooling, compute (GPU, TPU, other ASICs), network, structured cabling & passive infrastructure, storage, platform, operate. Canonical HTML: https://neuronplant.com/what-we-build Collection file: https://neuronplant.com/llms/system.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # Site The hall we occupy, the partner room, not a campus we pour. Floor load, kW, CDU, and the path in are constraints of that room. The hall owner supplies it. - Hall / colo cage we place into - Partner room or custom-system supplier - Rack planning in that room - Floor loading - U / weight / clearance - kW and heat the hall can present - Path in / meet-me Canonical: https://neuronplant.com/what-we-build#site # Power From the feed the hall presents to the rack PDU. Capacity, redundancy, and the path that holds under load. Rack PDUs: specify unmanaged, metered, or switched. Managed is not automatic. Utility interconnection is a hall-owner or custom-system-supplier job. - Hall-presented power - Capacity planning for the cell - UPS as the room allows - PDU / busway in the cell - Managed vs unmanaged PDU - Rack power distribution - Redundancy architecture Canonical: https://neuronplant.com/what-we-build#power # Cooling Heat is the constraint of the hall we place into. Liquid loops, CDUs, and rejection engineered to the water and rejection the hall owner supplies. - Liquid cooling - CDU design - Facility water requirements - Heat rejection - Chillers - Cooling redundancy Canonical: https://neuronplant.com/what-we-build#cooling # Compute Accelerators as an engineered cluster. GPU is the common case, not the only one. The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. NVIDIA DGX and GB-class remain how we size a GPU plant. Specs are NVIDIA's. The room we occupy, the loop, and the fabric are the design. We do not maintain a weekly chip catalog. - GPU (common case) - TPU - Other accelerators / ASICs - NVIDIA DGX / GB-class when the plant is GPU - Cluster architecture - Capacity planning Canonical: https://neuronplant.com/what-we-build#compute # Network East-West compute fabric first. Storage/data as its own plane when dedicated. North-South for users and applications. Management/OOB never as application North-South. - NVIDIA Quantum-X / Spectrum-X East-West compute - Storage / data fabric (dedicated when required) - North-South / front-end fabric - Management / OOB fabric - Converged vs dedicated compute and storage - 800G fabrics Canonical: https://neuronplant.com/what-we-build#network # Structured cabling & passive infrastructure Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. Skip it and the GPUs sit dark. - Fiber and MPO/MTP trunks (MMF / SMF, 400/800G optics) - Patch panels, cross-connects, rack-to-rack - Cable pathways and cable management - Copper and passive optical infrastructure - Separate trays for EW compute, storage/data, N-S, and OOB - Physical connectivity reserved for expansion - Labeling, IL / OTDR, polarity test, as-builts Canonical: https://neuronplant.com/cabling # Storage Datasets, checkpoints, RAG, and lifecycle, sized to the workload, not the brochure. - High-performance storage - Throughput / IOPS / metadata - AI datasets - Training storage - RAG sized from source data - Checkpointing - Backup and data lifecycle Canonical: https://neuronplant.com/what-we-build#storage # Platform The AI software layer after storage: serving, RAG services, NIM / runtime integration, and the path from data to workload. Accelerators without this platform are inventory. Large serving includes a KV and prefix cache hierarchy. The runtime that keeps it up is Operate. We do not write every model. - Platform architecture - Ingest and dataset path - Train / fine-tune vs RAG - Accelerator runtime placement - NVIDIA NIM profile sizing - KV / prefix cache (large serving) - Orchestration / MLOps interfaces - Enterprise application handover Canonical: https://neuronplant.com/applications # Operate The factory stays visible after handover. Docker ships the workload. Kubernetes is a 3+3 cell. Observe / monitor is part of that plane. Fluent Bit is the example shipper, not a dashboard catalog. - Docker images - Kubernetes cell (3 masters + 3 workers) - Masters may be VMs - CPU workers vs accelerator workers - Observe / monitor - Fluent Bit (example shipper) - Runbooks after handover Canonical: https://neuronplant.com/runtime # Rack PDU: managed vs unmanaged Managed PDU is a spec job, not a default. A PDU is how the rack gets power. Managed means the PDU has a management plane: metering (rack, phase, or outlet), network, often remote outlet control. Managed is not automatic. Specify the job: unmanaged distribution, metered, or switched at the outlet. This is not a vendor catalog. Managed PDU and unmanaged PDU are the durable words. ## Jobs - 01 Unmanaged distribution: Cheaper, fewer failure modes, still distributes power. A valid spec for dual-cord GPU and compute racks. - 02 Metered: See kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. Justified when the board or BMS does not already meter at this grain, or when the hall is multi-tenant or billed. - 03 Switched / outlet-level: Remote outlet control. On dual-cord GPU and compute racks this is a foot-gun: you do not power-cycle a training node from a PDU web UI. Specify it only when that job is real. ## When managed is justified - You need to see kW, phase balance, and headroom versus nameplate. Training loads are not nameplate. - Multi-tenant or billed halls. - The board or BMS does not already meter at this grain. ## When you do not need managed (or do not want outlet switching) - GPU and compute racks that are dual-corded. You do not power-cycle a training node from a PDU web UI. - EPMS or BMS already meters the feed. A metered PDU can be redundant. - The PDU management network is another plane to design, secure, and cable. Do not add it by default. - Unmanaged, or metered but not switched, is a valid spec. Cheaper, fewer failure modes, still distributes power. Pages: /what-we-build#pdu and /capabilities#pdu. Power plant layer. Not a SKU list. Canonical: https://neuronplant.com/capabilities#pdu # Facility as architecture of the hall we occupy The hall is architecture. Rack space is not power or cooling. A hall we occupy is a set of coupled limits: U, weight, floor load, normal and peak kW, A/B, PDU, heat, clearance, and pathways. Empty U with a dead feed is not capacity. The hall owner supplies the room. NeuronPlant engineers the GPU / TPU / ASIC cell to fit, or matches another hall. We do not pour the slab. - U / rack space: Necessary and almost never sufficient. Count the accelerators, CDU manifolds, patch, and service U. Then ask whether power and heat still fit. - Weight: Liquid-cooled GPU racks and full storage frames are plant loads. If the slab or raised floor cannot take them, the SKU does not ship, or it ships and sits. - Floor load: kN/m², not a vibe. Point loads under NVL-class racks and CDUs. Path in from the dock matters as much as the pad. - Normal and peak kW: Nameplate is not the bill. Training peaks. Inference is different. Design the feed, UPS, and breaker to the number that shows up, A/B where the architecture requires it. - A/B: Dual cord is a drawing: two PDUs, two paths, maintenance bypass. Dual cord on a single upstream is a sticker. - PDU: Circuit count, phase, and whether managed is justified. See Power / PDU. Reserve circuits for the additive cell or admit the row is full. - Heat: Rejection path, not a CRAH brochure. Air until physics fails. Then water, CDU, and isolation. Peak kW is peak heat. - Clearance: Service, manifolds, and a sled that can come out without tensioning fiber. A packed hot aisle is not density. It is an unmaintainable plant. - Pathways: Fiber baskets, power, liquid. Overhead versus underfloor. The next trunk needs a reserved path. See Cabling / tomorrow. Pages: /capabilities#facility, /what-we-build#site. Expansion gate: /expansion#gates. Canonical: https://neuronplant.com/capabilities#facility