Skip to main content
NP-RT/Runtime / Operate

The factory runs as a fleet of images, not as a snowflake hall.

After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane.

Durable names. Not a SKU catalog.

DWG-RT-01 · laying out

01 / Why Docker

Why Docker

The workload ships as an image.

Serving, ingest, RAG, NIM runtime, plant services. The hall is not a unique snowflake OS. Same artifact from lab to factory.

02 / Why Kubernetes

Why Kubernetes

Schedule, restart, and place those containers across the plant.

Accelerator nodes, CPU nodes, and plant services in one control plane. Not because cloud native is a slogan. An AI factory is a fleet. A fleet that is ssh-and-excel will not stay up.

NP-RT-03/What Kubernetes requires

A starting cell: three masters and three workers. That is a plant constraint, not a certification syllabus.

Three masters

Three masters so etcd and the API stay up if one dies. That is why 3, not 1.

Three workers

Three workers as a starting factory cell. The cell can grow. Accelerator workers are not the same machine class as the masters.

Masters may be VMs

Masters may be virtual machines. Workers that hold accelerators are usually bare metal. The plant does not require three physical master servers.

Plant jobs, not a SKU list

Network, storage classes, and an accelerator device plugin are jobs on the plane. They are not a product catalog we refresh every week.

NP-RT-05/Kubernetes as production infrastructure

Not K8s plus vLLM as a slogan. A factory control plane with CPU workers, accelerator workers, and the platform services that must not sit on the expensive nodes.

Control plane

API, etcd, schedulers. Masters may be VMs. This plane stays up if a GPU node dies.

CPU workers

Ingest, parse, RAG services, registries, identity, plant jobs. Sized as a cell, grown as a gate. See /expansion.

Accelerator workers

GPU is the common case. TPU and other ASICs when that is the job. Bare metal. Device plugin. Not a place to park Prometheus.

Registry

The images the factory actually runs. On the management plane, not on an HBM node.

Secrets

A named store and a rotation path. Not files on the GPU worker.

Identity

Who may schedule, who may retrieve, who may delete. RAG ACL starts here and has to reach the index.

RAG services

Parse, embed jobs, retrievers. CPU first. Accelerators only for the embed/infer step that needs them.

Keep platform off accelerator nodes

Do not park registry, logging, identity, or the control plane on expensive accelerator workers. That is how a fleet starves itself.

Platform growth is a gate, not a slogan. Expansion / platform.

04 / Observability

Observability

Logs, metrics, and traces are part of the plane.

Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards.

In scope

  • Docker as the image that moves serving, ingest, RAG, NIM, and plant services from lab to factory
  • Kubernetes as the control plane that schedules that fleet across accelerator, CPU, and plant nodes
  • A starting cell of three masters and three workers, with masters allowed to be VMs
  • Observability as part of the plane, with Fluent Bit as the example log shipper
  • Production Kubernetes: control plane, CPU workers vs accelerator workers, registry, secrets, identity, RAG services
  • Platform services kept off expensive accelerator nodes

Not in scope

  • A CNCF certification syllabus or weekly operator-version pin list
  • A CNI, service-mesh, or dashboard catalog we would have to refresh every week
  • Requiring three physical master servers when VMs will hold quorum
  • Treating ssh-and-excel as the runbook for an accelerator hall
  • K8s plus vLLM as the whole platform story
  • Named TTFT / ITL numbers as a NeuronPlant product
Sits on the platform after accelerator runtime placement. Applications / Platform · What We Build.

NP / Intake

Tell us what you need to achieve with AI.

One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.

Start a Project →