Skip to main content
NP-VRAM/GPU memory

How much VRAM does this context actually take?

Weights, KV cache, and concurrency, from the model's published layer and KV-head counts. A planning estimate for a NeuronPlant delivery. Not an exact runtime measurement.

Mode
Model

36 layers · 32 Q heads · 8 KV heads · head 128 · 8.2B

Weight precision
KV cache precision
Hardware

96 GB GDDR7 ECC · 600 W. Workstation Edition. 1,792 GB/s. Published MIG: 4×24 GB, 2×48 GB, or 1×96 GB. Slices do not share weights.

Context length
Concurrent sequences
Safety margin

NP / Intake

Tell us what you need to achieve with AI.

One accountable delivery partner from AI requirement to a delivered factory: site, power, cooling, compute, fabric, structured cabling & passive infrastructure, storage, platform, procurement, and handover. You do not need a vendor list or a hall name first. The hall is a supplier.

Start a Project →