# NeuronPlant / Cited models NVIDIA NIM support-matrix profiles used to size nodes vs racks. Canonical HTML: https://neuronplant.com/models Collection file: https://neuronplant.com/llms/model.txt NeuronPlant is not an NVIDIA Cloud Partner. NVOnline packs are named, not reprinted. Customer names and customer specifications publish only after written approval. # Llama 3.3 Nemotron Super 49B v1.5 Use: Enterprise reasoning / chat NVIDIA NIM publishes BF16, FP8 and NVFP4 profiles at TP 1, 2, 4 and 8, with and without LoRA. Fits a single NVIDIA DGX B200 or NVIDIA DGX B300 for many FP8/NVFP4 profiles. TP8 is the full 8-GPU node. Size KV cache and concurrency separately. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/models # Nemotron 3 Super 120B-A12B Use: Larger MoE / reasoning NVIDIA NIM lists BF16, FP8 and NVFP4 across TP 1-8, base and LoRA. Needs more GPU memory than a 49B. On Blackwell Ultra (NVIDIA DGX B300) dense FP4/NVFP4 is the practical serving path. Training or distillation belongs on a coupled fabric, not on the inference node. NVIDIA source: https://docs.nvidia.com/nim/large-language-models/latest/support-matrix.html Canonical: https://neuronplant.com/models # Trillion-parameter / MoE class Use: Frontier inference NVIDIA describes GB200 NVL72 as a 72-GPU NVLink domain for real-time trillion-parameter LLM inference and MoE. This is a rack (or many racks), liquid loop, and scale-out fabric problem. Do not size it as nine 8-GPU servers on a 100G ToR. NVIDIA source: https://www.nvidia.com/en-us/data-center/gb200-nvl72/ Canonical: https://neuronplant.com/models