Skip to main content
NP-MDL/Models

NVIDIA references size a GPU plant.

NeuronPlant does not publish its own model cards. The names, profiles and system claims below are taken from NVIDIA NIM and NVIDIA DGX / GB200 public documentation. We use them to size node vs rack, fabric, storage, and the inference serving plane when delivering a GPU plant.

NP-MDL / Models drive the factory

Start from the model, then the rack.

Profiles and product claims below are taken from NVIDIA public documentation. NeuronPlant uses them as engineering inputs, not as our own benchmarks.

Enterprise reasoning / chat

Llama 3.3 Nemotron Super 49B v1.5

NVIDIA NIM publishes BF16, FP8 and NVFP4 profiles at TP 1, 2, 4 and 8, with and without LoRA.

Fits a single NVIDIA DGX B200 or NVIDIA DGX B300 for many FP8/NVFP4 profiles. TP8 is the full 8-GPU node. Size KV cache and concurrency separately.

Larger MoE / reasoning

Nemotron 3 Super 120B-A12B

NVIDIA NIM lists BF16, FP8 and NVFP4 across TP 1-8, base and LoRA.

Needs more GPU memory than a 49B. On Blackwell Ultra (NVIDIA DGX B300) dense FP4/NVFP4 is the practical serving path. Training or distillation belongs on a coupled fabric, not on the inference node.

Frontier inference

Trillion-parameter / MoE class

NVIDIA describes GB200 NVL72 as a 72-GPU NVLink domain for real-time trillion-parameter LLM inference and MoE.

This is a rack (or many racks), liquid loop, and scale-out fabric problem. Do not size it as nine 8-GPU servers on a 100G ToR.

Sources: NVIDIA NIM for LLMs support matrix; NVIDIA GB200 NVL72.

The compute layer is accelerators. GPU is the common case, not the only one. TPU is a named family buyers already specify. The third class is other accelerators and ASICs built to run a model or a graph fast. We do not maintain a weekly chip catalog. NIM is the cited inference runtime on the AI platform, not a NeuronPlant product. Profile sizing belongs with the plant: local NVMe, planes, and concurrency. See the platform layer.

NVIDIA platforms used to size a GPU plant

Node · 8x NVIDIA Blackwell

NVIDIA DGX B200

Unified develop-to-deploy node. NVIDIA lists 1,440 GB GPU memory and up to 400 Gb/s ConnectX-7 / BlueField-3 ports.

Node · 8x NVIDIA Blackwell Ultra

NVIDIA DGX B300

NVIDIA positions NVIDIA DGX B300 for reasoning / inference density. Docs list 8x B300 SXM, ConnectX-8 up to 800 Gb/s, BlueField-3 for storage and management.

Rack · 72 Blackwell + 36 Grace

NVIDIA DGX GB200 NVL72

Liquid-cooled rack-scale NVLink domain. NVIDIA describes the GB200 NVL72 architecture as a 72-GPU NVLink domain for large-model inference and MoE.

Rack · 72 Blackwell Ultra + 36 Grace

NVIDIA DGX GB300 NVL72

NVIDIA DGX rack-scale system. Published networking includes ConnectX-8 at 800 Gb/s InfiniBand and NVLink Switch System.

NP / Intake

Tell us what you need to achieve with AI.

One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.

Start a Project →