# NeuronPlant / Factory templates Private AI, training, AI-as-a-Service, sovereign, and large-scale factories NeuronPlant can deliver. Canonical HTML: https://neuronplant.com/ai-factories Collection file: https://neuronplant.com/llms/factory.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # Enterprise Private AI Factory Private enterprise AI / RAG / inference. A controlled environment for enterprise models, retrieval, and inference, inside the organization's security and data boundary. ## Pipeline - Data Sources - Ingestion - Object / File Storage - Vector / RAG Layer - Compute - Inference Platform - Enterprise Applications - Operate ## Infrastructure - Site: Inside the organization's security and data boundary. - Power: Rack-level redundancy, typically 20-40 kW/rack - Cooling: Air or targeted liquid depending on density - Compute: Accelerators for NIM / RAG serving. GPU plant: NVIDIA DGX B200 or NVIDIA DGX B300, or HGX. TPU and other ASICs when that is the job. - Network: Spectrum-X East-West compute fabric, dedicated storage/data path, separate North-South front-end - Structured cabling: Labeled LC/MPO plant. Management/OOB on its own patch field. Factory-tested trunks. Polarity map before first optic. - Storage: Object for documents + NVMe or memory vector index. Model weights on local NVMe. - Platform: Sources → ingest → RAG / inference → enterprise applications. - Operate: Docker / Kubernetes cell. Observe as part of the plane. ## Considerations - Dataset size and retention - Growth of concurrent sessions - Availability target - Security and data residency - Latency to enterprise applications - Data lifecycle and audit - Whether the existing hall plant can take 400G without a recable - How the serving plane is observed after handover Canonical: https://neuronplant.com/ai-factories/private-ai # High-Performance Training Factory High-performance accelerator training infrastructure. A tightly coupled training cluster: fabric, storage bandwidth, and cooling sized to the job, not the idle rack. GPU is the common case. TPU and other ASICs when that is the domain. ## Pipeline - Datasets - Staging / Parallel FS - Training fabric - Checkpoint Store - Scheduler - Experiment Platform - Operate ## Infrastructure - Site: Dedicated training hall we occupy. Floor, water, and power are gates of that room before the first node. - Power: High-density feeds, often 40-120 kW/rack - Cooling: Direct liquid cooling and CDU-backed rejection - Compute: Accelerators in one training domain. GPU plant: NVIDIA DGX B200 or NVIDIA DGX B300 nodes, or NVIDIA DGX GB200 NVL72 or NVIDIA DGX GB300 NVL72 racks. TPU and other ASICs sized the same way, not as a catalog. - Network: Quantum InfiniBand East-West compute fabric; dedicated storage/data fabric. Cited SuperPOD jobs on Capabilities. - Structured cabling: Rail-aligned MPO trunks. Type-B polarity as a drawing, not a field guess. East-West compute, storage/data, and management/OOB on separate trays. IL and polarity records before first job. - Storage: HPS parallel FS sized to NVIDIA SuperPOD guidance (>40 GB/s per NVIDIA DGX B300 node; >80 GB/s on XDR RA) plus a checkpoint burst target - Platform: Datasets → staging → train fabric → checkpoints → scheduler. - Operate: 3+3 Kubernetes cell. Observe the fabric and the job. ## Considerations - Model scale and parallelism strategy - Job mix vs dedicated cluster - Checkpoint frequency and recovery - Fabric congestion under all-reduce - Structured cabling: polarity, slack, and whether the hall can be tested - Facility power and water lead times - Expansion without re-architecture - Who watches the fabric and the job after first run Canonical: https://neuronplant.com/ai-factories/training # Multi-Tenant Compute Factory Multi-tenant accelerator infrastructure designed for service providers. A factory built to sell capacity: isolation, metering, rapid provisioning, and a hall-partner model that can grow in cells. ## Pipeline - Onboarding - Tenancy / Isolation - Accelerator pools - Shared Fabric - Platform Control Plane - Billing / Observability - Operate ## Infrastructure - Site: Colo or owned hall from a hall partner, grown in cells. Not a NeuronPlant campus. - Power: Cell-based capacity, N+1 at hall and rack - Cooling: Liquid-ready rows for density growth - Compute: Pooled accelerators. GPU nodes with MIG / tenant slicing where applicable. - Network: Segmented Spectrum-X or Quantum fabrics: East-West compute isolated from tenant North-South, noisy-neighbor controls - Structured cabling: Per-plane trays so a tenant storage change cannot yank an East-West compute trunk. Labels that survive a night shift. - Storage: Per-tenant namespaces on a shared high-performance backend - Platform: Onboarding → tenancy → accelerator pools → control plane → customer workloads. - Operate: Shared control plane, per-tenant isolation, observe as a billed plane. ## Considerations - Tenant isolation model - Oversubscription policy - Support and SLA boundaries - Chargeback and telemetry - Cable plant isolation between tenants and between planes - Cooling headroom for densification - Sovereign or regional cells - Tenant telemetry without a dashboard shopping list Canonical: https://neuronplant.com/ai-factories/ai-as-a-service # Sovereign AI Factory Private, controlled AI infrastructure for regulated organizations. Infrastructure that can be operated, inspected, and expanded under a defined jurisdiction: power to platform, not only the model weights. ## Pipeline - National / Org Data - Controlled Ingest - Sovereign Storage - Private compute domain - Governed Platform - Authorized Workloads - Operate ## Infrastructure - Site: In-jurisdiction hall from a partner or custom-system supplier. Inspectable plant and operational control. - Power: Facility-level independence and documented failover - Cooling: Heat rejection the hall owner or custom-system supplier presents, with operational control in-jurisdiction. - Compute: Dedicated accelerator domains, no shared cloud control plane - Network: Air-gapped or tightly gated interconnects - Structured cabling: Inspectable plant: labeled, tested, as-built. No undocumented jumpers on the air-gap. - Storage: In-country datasets, encryption and key custody - Platform: National / org data → controlled ingest → governed platform → authorized workloads. - Operate: In-jurisdiction runtime cell. Observe stays inside the boundary. ## Considerations - Data residency and custody - Personnel and operational control - Supply chain of compute - Auditability of the stack - Independent support model - Long-term expansion path - Observe stays inside the jurisdiction Canonical: https://neuronplant.com/ai-factories/sovereign-ai # Large-Scale AI Factory Large-scale accelerator environment with liquid cooling and high-speed networking. A hall-scale accelerator program delivered with a hall partner or custom-system supplier: megawatts and liquid as hall constraints, 800G fabrics and the GPU / TPU / ASIC plant as NeuronPlant's job. Not NeuronPlant as the data-center contractor. ## Pipeline - Datasets / ingest - Hall-scale HPS - Accelerator supercluster - Checkpoint / archive - AI platform - Production workloads - Operate ## Infrastructure - Site: MW-class hall from a partner or custom-system supplier. Utility and water are gates of that room, not a NeuronPlant civil works package. - Power: MW-class, staged to utility and on-site distribution - Cooling: Facility water, CDUs, rear-door or direct-to-chip, N+1 rejection - Compute: Accelerator halls at multi-pod scale. GPU plant: NVIDIA DGX GB200 NVL72 / NVIDIA DGX GB300 NVL72. TPU and other ASICs when that is the job. - Network: Quantum-X or Spectrum-X East-West scale-out, dedicated storage/data fabric, North-South for users, management/OOB independent - Structured cabling: SMF/MPO campus plant inside the CIN 500 m class from the public NVIDIA DSX overview. Overhead baskets. Trays grouped by plane: East-West compute, storage/data, North-South, management/OOB. OTDR on backbone. As-builts before first power-on. - Storage: Multi-tier: scratch, dataset, checkpoint, archive - Platform: Datasets → hall HPS → supercluster → checkpoints → production workloads. - Operate: Cell-by-cell Kubernetes. Observe before first training run, not after. ## Considerations - Utility interconnection timeline (hall owner or custom-system supplier) - Water and heat-rejection path - Phased pod delivery - Logistics of high-density racks - Fiber plant reach, polarity, and test records at hall scale - Operations before first training run - Cell-by-cell expansion - Observe before first training run Canonical: https://neuronplant.com/ai-factories/large-scale