Three masters
Three masters so etcd and the API stay up if one dies. That is why 3, not 1.
After the accelerators exist, the workload still has to ship, schedule, and stay up. Docker is the artifact. Kubernetes is the plant's control plane. That plane rides management, not application North-South. A starting cell is three masters and three workers. Observability is part of that plane.
Durable names. Not a SKU catalog.
DWG-RT-01 · laying out
01 / Why Docker
The workload ships as an image.
Serving, ingest, RAG, NIM runtime, plant services. The hall is not a unique snowflake OS. Same artifact from lab to factory.
02 / Why Kubernetes
Schedule, restart, and place those containers across the plant.
Accelerator nodes, CPU nodes, and plant services in one control plane. Not because cloud native is a slogan. An AI factory is a fleet. A fleet that is ssh-and-excel will not stay up.
Three masters
Three masters so etcd and the API stay up if one dies. That is why 3, not 1.
Three workers
Three workers as a starting factory cell. The cell can grow. Accelerator workers are not the same machine class as the masters.
Masters may be VMs
Masters may be virtual machines. Workers that hold accelerators are usually bare metal. The plant does not require three physical master servers.
Plant jobs, not a SKU list
Network, storage classes, and an accelerator device plugin are jobs on the plane. They are not a product catalog we refresh every week.
Control plane
API, etcd, schedulers. Masters may be VMs. This plane stays up if a GPU node dies.
CPU workers
Ingest, parse, RAG services, registries, identity, plant jobs. Sized as a cell, grown as a gate. See /expansion.
Accelerator workers
GPU is the common case. TPU and other ASICs when that is the job. Bare metal. Device plugin. Not a place to park Prometheus.
Registry
The images the factory actually runs. On the management plane, not on an HBM node.
Secrets
A named store and a rotation path. Not files on the GPU worker.
Identity
Who may schedule, who may retrieve, who may delete. RAG ACL starts here and has to reach the index.
RAG services
Parse, embed jobs, retrievers. CPU first. Accelerators only for the embed/infer step that needs them.
Keep platform off accelerator nodes
Do not park registry, logging, identity, or the control plane on expensive accelerator workers. That is how a fleet starves itself.
Platform growth is a gate, not a slogan. Expansion / platform.
04 / Observability
Logs, metrics, and traces are part of the plane.
Fluent Bit is the example log shipper. A DaemonSet at the edge of every node, forwarding to the plant's log store. That is the note. Not a weekly shopping list of dashboards.
In scope
Not in scope
NP / Intake
One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.