# NeuronPlant / Site routes Important pages and what they contain. Canonical HTML: https://neuronplant.com/ Collection file: https://neuronplant.com/llms/route.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. NeuronPlant. We Build AI Factories. One Partner. Design Authority and Delivery. The Entire AI Factory. You bring the AI requirement. NeuronPlant delivers the factory. The customer comes to NeuronPlant. Complexity behind the scenes belongs to NeuronPlant, not to the customer. The hall is a supplier. Not a data-center contractor. NeuronPlant is the design authority and the single accountable delivery partner for an AI factory. You bring the AI requirement. NeuronPlant engineers the architecture, can purchase the required equipment against that architecture, coordinates suppliers, integrates, validates, and hands over the factory. We do not build the data center. The hall operator supplies and operates the room. Homepage is the buying story: hero, problem, one journey (Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover), what NeuronPlant owns, canonical nine-layer plant. Facility fit and architecture come before procurement. NeuronPlant can purchase the required equipment against that architecture. It is not a box reseller taking a customer shopping list. The customer calls NeuronPlant. They do not need a BOM, a hall name, or a vendor list. What NeuronPlant owns: hall as supplier, nine-layer plant, and the delivery. Canonical plant: Site → Power → Cooling → Compute → Network → Structured cabling → Storage → Platform → Operate. AI Factory System View is on /what-we-build#system-view as an unnumbered dependency stack, not a second layer model. GPU COMPUTE on the hero is labeled ACCELERATED COMPUTE (GPU · TPU · ASIC). KV-cache, RAG lifecycle, Fluent Bit, PDU outlet switching, Kubernetes worker placement, and storage-sizing theory stay off the homepage. NeuronPlant does not construct the data center. The hall operator supplies and operates the room. Island actions: Start a Project and Become a Partner. Portal is the customer's project environment preview, not a hall marketplace, and not the product. MCP, llms.txt, and agent discovery live in the footer. Canonical: https://neuronplant.com/ Plant layers: site (the hall we occupy / partner room, not a campus NP pours), power (managed vs unmanaged PDU is a spec job), cooling, compute (accelerators: GPU, TPU, and other ASICs; NVIDIA DGX / GB-class when the plant is GPU), network, structured cabling & passive infrastructure, storage, platform (serving, RAG services, NIM / runtime integration; the data path is still a pipeline), operate (Docker, Kubernetes 3+3 cell, observe / monitor, Fluent Bit as the example shipper). Nine layers is the plant NeuronPlant delivers as one project. The AI Factory System View lives at /what-we-build#system-view as an unnumbered dependency stack (utility, hall, plant, workload, applications), not a second numbered layer model. Commercial journey is Requirement → Facility Fit → Architecture → Procurement → Coordination → Deployment → Validation → Handover. Expansion / Day-1 (/expansion) is additive vs transition, gates, and min vs future-ready. Canonical: https://neuronplant.com/what-we-build Capabilities are organized by accelerator-plant layer. Facility (U, weight, kW, A/B, heat, pathways) sits with Site as constraints of the hall we occupy. The hall owner supplies the room. Rack PDUs (managed vs unmanaged) sit with Power. Fabrics: four planes, not four equal jobs. East-West compute, storage/data when dedicated, North-South front-end, management/OOB. Compute vs storage is a converged/dedicated decision. Networking is more than port count. Storage: throughput, IOPS, metadata, checkpoints, RAG from source data, protection. Operate / monitor sits with the runtime cell. NVIDIA switch families, storage profiles, and cited NIM models live here. NVIDIA DSX sits with AI Factories, not here. NeuronPlant is not the data-center general contractor. Canonical: https://neuronplant.com/capabilities Each factory is a template for a NeuronPlant delivery: data pipeline, infrastructure by layer, and design considerations. Site in each template is the hall we occupy. NVIDIA DSX is how we read campus-scale language from public docs when a hall partner or customer program needs it, not a campus NeuronPlant builds. Canonical: https://neuronplant.com/ai-factories NeuronPlant does not publish original model cards. NVIDIA NIM and DGX/GB references size a GPU plant. Compute is accelerators: GPU, TPU, and other ASICs. We do not maintain a weekly chip catalog. Canonical: https://neuronplant.com/models Map of grid, BESS, CUB, dry coolers, CDU gallery, compute hall, CIN, restated so a hall partner or customer program can speak that language. NVIDIA's 157-acre example is NVIDIA's published table, not a NeuronPlant delivery. Not an NVOnline reprint. Canonical: https://neuronplant.com/dsx Fiber, MPO/MTP trunks, structured cabling, patch panels and cross-connects, rack-to-rack connectivity, cable pathways and cable management, copper, passive optical infrastructure, and physical connectivity prepared for future expansion. The layer that lands the fabric. Skip it and the GPUs sit dark. Canonical: https://neuronplant.com/cabling # NeuronPlant expansion / Day-1 Growth is an architecture evolution, not a node count. Adding capacity inside a designed envelope is additive. Changing the fabric, the cooling loop, the storage machine, or the hall is a transition. Day-1 names those points. It does not promise unlimited scale. We do not sell a slogan that scales two DGX nodes to eight. We say how far Day-1 grows without a redesign, and what measurement opens the next investment. Plant sequence: Site → Power → Cooling → Compute → Network → Structured cabling & passive infrastructure → Storage → Platform → Operate. ## Additive vs transition - Additive expansion: More of the same cell inside a designed envelope: another accelerator node on the same fabric, another PDU circuit in the reserved row, another NVMe shelf on the same storage machine. The drawings already have a place for it. - Architecture transition: The envelope is full. East-West compute changes class. Cooling moves from air to liquid. Storage splits. The hall we occupy cannot take the next kW. That is a new architecture, or a different hall supplier, not a purchase order for two more nodes. NeuronPlant does not pour a new campus to keep the slogan. ## How far Day-1 grows without a redesign Day-1 is a plant with named headroom, not a rack that happens to work. Write the first transition before the first PO. Then the additive path is honest. - Accelerator count the fabric, PDU, and loop can take without a recable or a new CDU. - Storage throughput and metadata the chosen machine can take before it becomes a different product. - Pathway, MPO, and switch-position reserve that the next leaf can land on. - Platform cell (CPU workers, registry, identity) that does not have to move onto accelerator nodes as the fleet grows. - Facility: floor load, peak kW, heat, clearance, and the next cage or row of the hall we occupy, or a match to another hall. Not a campus NeuronPlant pours. ## Expansion gates What measurement triggers the next investment. Not unlimited scale. - Compute: HBM, concurrency, or job queue time. The next node still fits the fabric, PDU, and loop, or it does not. Measurement: Utilisation and wait, not a wish for more GPUs. GPU is the common case. TPU and other ASICs use the same gate. - Storage: Throughput, IOPS, metadata, or protection window. Capacity full is often the last signal, not the first. Measurement: Bytes moved and files created under the real job, against the envelope written at Day-1. - Fabric: Oversubscription, optics class, leaf positions, or a plane that can no longer stay isolated. Measurement: Congestion and blast radius on East-West compute, storage/data, North-South, and management/OOB. Port count is not the gate. - Platform: Control-plane load, registry, identity, RAG services, or CPU workers saturating while accelerators wait. Measurement: Whether production Kubernetes still has a CPU plane. Do not grow the fleet by parking services on GPU nodes. - Facility: Normal versus peak kW, floor load, heat, A/B, PDU circuits, clearance, pathways. Rack U left is not power left. Measurement: What the hall we occupy can take. If it cannot, match a hall partner or custom-system supplier. NeuronPlant does not become the DC. ## Two Day-1 strategies - Minimum Day-1. CapEx: Lower first invoice. Buy what runs the first job. Disruption: The first transition arrives sooner: recable, new loop, new storage machine, or a hall that cannot take the next cell. When: The workload is bounded, the hall is a hard cage, and the customer accepts a written stop. - Future-ready Day-1. CapEx: More first money in trunks, pathways, PDU reserve, switch positions, CPU workers, and cooling that the next cell can join. Disruption: Additive growth stays additive longer. The transition, when it comes, is a named gate, not a surprise outage. When: The fleet will grow, the hall can take a foundation, and disruption costs more than copper and water now. ## The money path stays visible. Minimum Day-1 and future-ready Day-1 are both honest if the customer sees the same CapEx, disruption, and next-gate cost we see. Integrity before a steered BOM still applies: we do not hide the cheaper first invoice that forces a redesign. Unlimited scale is not an option on the form. ## Products are not the architecture. Coupled constraints are. An AI factory is not DGX plus a switch plus storage plus Kubernetes. It is constraints that couple. Engineering those dependencies happens before procurement. - Plant chain: Model → accelerator → network → storage → power → cooling → rack → facility → ops → cost. Change the model class and the accelerator, fabric, loop, and hall move with it. GPU is the common case. TPU and other ASICs still travel this chain. - Data chain: Data → ingest → RAG → inference → KV → capacity. Change the corpus, the embedding, or the session length and the storage machine, the serving plane, and the KV hierarchy move with it. A vector SKU does not size that chain. Canonical page: /expansion. Intake money path: /start-a-project. Homepage carries one sentence that Day-1 names its transition points. It does not reprint this page. Canonical: https://neuronplant.com/expansion Accelerators without a platform are inventory. The factory includes serving, RAG services, NIM / runtime integration, and the path from data to workload. Canonical: https://neuronplant.com/applications # NeuronPlant runtime / operate The factory runs as a fleet of images, not as a snowflake hall. After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane. ## Why Docker The workload ships as an image. Serving, ingest, RAG, NIM runtime, plant services. The hall is not a unique snowflake OS. Same artifact from lab to factory. ## Why Kubernetes Schedule, restart, and place those containers across the plant. Accelerator nodes, CPU nodes, and plant services in one control plane. Not because cloud native is a slogan. An AI factory is a fleet. A fleet that is ssh-and-excel will not stay up. ## What Kubernetes requires A starting cell: three masters and three workers. That is a plant constraint, not a certification syllabus. - Three masters: Three masters so etcd and the API stay up if one dies. That is why 3, not 1. - Three workers: Three workers as a starting factory cell. The cell can grow. Accelerator workers are not the same machine class as the masters. - Masters may be VMs: Masters may be virtual machines. Workers that hold accelerators are usually bare metal. The plant does not require three physical master servers. - Plant jobs, not a SKU list: Network, storage classes, and an accelerator device plugin are jobs on the plane. They are not a product catalog we refresh every week. ## Observability Logs, metrics, and traces are part of the plane. Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards. ## Kubernetes as production infrastructure Not K8s plus vLLM as a slogan. A factory control plane with CPU workers, accelerator workers, and the platform services that must not sit on the expensive nodes. - Control plane: API, etcd, schedulers. Masters may be VMs. This plane stays up if a GPU node dies. - CPU workers: Ingest, parse, RAG services, registries, identity, plant jobs. Sized as a cell, grown as a gate. See /expansion. - Accelerator workers: GPU is the common case. TPU and other ASICs when that is the job. Bare metal. Device plugin. Not a place to park Prometheus. - Registry: The images the factory actually runs. On the management plane, not on an HBM node. - Secrets: A named store and a rotation path. Not files on the GPU worker. - Identity: Who may schedule, who may retrieve, who may delete. RAG ACL starts here and has to reach the index. - RAG services: Parse, embed jobs, retrievers. CPU first. Accelerators only for the embed/infer step that needs them. - Keep platform off accelerator nodes: Do not park registry, logging, identity, or the control plane on expensive accelerator workers. That is how a fleet starves itself. ## In scope - Docker as the image that moves serving, ingest, RAG, NIM, and plant services from lab to factory - Kubernetes as the control plane that schedules that fleet across accelerator, CPU, and plant nodes - A starting cell of three masters and three workers, with masters allowed to be VMs - Observability as part of the plane, with Fluent Bit as the example log shipper - Production Kubernetes: control plane, CPU workers vs accelerator workers, registry, secrets, identity, RAG services - Platform services kept off expensive accelerator nodes ## Out of scope - A CNCF certification syllabus or weekly operator-version pin list - A CNI, service-mesh, or dashboard catalog we would have to refresh every week - Requiring three physical master servers when VMs will hold quorum - Treating ssh-and-excel as the runbook for an accelerator hall - K8s plus vLLM as the whole platform story - Named TTFT / ITL numbers as a NeuronPlant product Canonical page: /runtime. Also on /what-we-build#operate and /applications#runtime. Canonical: https://neuronplant.com/runtime Records stay anonymized until a customer authorizes a name in writing. Canonical: https://neuronplant.com/projects Same approval gate as customer specifications in the published factory record. Canonical: https://neuronplant.com/customers NVL72 deployment, power and cooling, Ethernet vs InfiniBand, structured cabling failures, steered RFPs. Canonical: https://neuronplant.com/insights NeuronPlant is the design authority and the single accountable delivery partner for the GPU, TPU, and ASIC plant: compute, fabrics, structured cabling and passive infrastructure, AI storage, platform, runtime, KV, RAG. NeuronPlant engineers the architecture, can purchase the required equipment against that architecture, coordinates suppliers, and delivers. The hall is a supplier: an existing colo or a custom-system builder. NeuronPlant does not construct data centers. The customer calls NeuronPlant. They do not assemble the vendor list. Integrity before margin on a steered BOM. Purchasing equipment does not mean hiding the money path. The customer sees the same constraints, options, and money path we see. No games. No tricks. Company pages: About, About / Integrity (#integrity), Partners, Become a Partner, Customers. Canonical: https://neuronplant.com/about Partners are split into Technology and Delivery. Technology (the stack), by layer: Site (cage / floor / path-in interfaces of the hall we occupy, not a campus NP constructs), Power, Cooling, Compute (NVIDIA, Lenovo, OEM accelerator systems: GPU, TPU, and other ASICs), Network (NVIDIA Quantum InfiniBand, NVIDIA Spectrum-X Ethernet), Storage (VAST, DDN, parallel/object). NVIDIA and Lenovo are listed as technology companies in the stack, not as NeuronPlant partners. No NVIDIA or Lenovo logos. Delivery (who ships it): Hall (existing hall / colo operators; custom-system hall suppliers. NeuronPlant is not the colo and does not pour the slab. /portal is the customer's project record, not a hall marketplace), MyIT Cyber (Integration, https://myitcyber.com) and Dayo-Tech (Compute-delivery, https://www.dayo-tech.com). Delivery companies sit behind NeuronPlant. Dayo-Tech's public NVIDIA Elite Partner claim is Dayo-Tech's, not NeuronPlant's. Named Delivery companies appear only when authorized. Partner registration is /become-a-partner with required type Technology or Delivery. Hall is a plant layer on that form. Canonical: https://neuronplant.com/partners POST /api/partner. Required: company, type (Technology or Delivery), layer (site/hall/power/cooling/compute/network/storage/integration/software), supply, geography, nvidiaRelation, name, email. Hall and colo operators belong as Delivery / hall. Success returns a ticket id. Header island has an outline Become a Partner button to this intake (Partner on tight widths), next to Start a Project. Same action in the mobile menu. Not labeled BETA. Company mega-menu still lists Partners and Become a Partner. Footer keeps a quiet Partner line. Canonical: https://neuronplant.com/become-a-partner NeuronPlant runs the delivery as one customer project record: status, architecture, facility and hall, suppliers, procurement, milestones, documentation, and handover. /portal is a commissioning preview of that environment. It is not a marketplace connecting halls with customers. Submit does not log anyone in and does not invent data. Until handover, the record stays with NeuronPlant. A future login would use the same project ticket issued at /start-a-project (NP-). Partner tickets stay on the NP-P- series. Not a header collection. Linked from Start a Project, About, and the footer. NeuronPlant does not operate the colo. Canonical: https://neuronplant.com/portal Public Streamable HTTP endpoint at /api/mcp. Read-only. No auth. Canonical: https://neuronplant.com/mcp Short map of routes and /llms/*.txt collection files. Full concatenated corpus: /llms-full.txt, organized by collection. Canonical: https://neuronplant.com/llms.txt