How to read 12 GB, 16 GB, and 24 GB cards for local LLMs

Treat marketing GB as GiB, subtract a reserve, then compare a range — not weights versus the box label.

Updated 2026-09-19. Figures use calculator 0.1.0 unless the article says otherwise.

NVIDIA lists GeForce RTX 4070 as 12 GB, RTX 4080 SUPER as 16 GB, and RTX 4090 as 24 GB GDDR6X. This catalog stores those as 12, 16, and 24 GiB of discrete VRAM, because GDDR capacities are binary modules. The assumption is written on each hardware page. Host RAM is not added to those numbers.

Default reserves in this dataset: 1 GiB on 12–16 GB cards, 2 GiB on the 24 GB 4090. If you already typed a net budget, the reserve is not subtracted again.

Worked screens, calculator 0.1.0, 8,192 tokens, 1 sequence, FP16 KV, llama.cpp typical extras unless noted:

Always compare the range to the net budget. Comparing only the GGUF size to “24GB” on the box is the usual way to get a surprise OOM.

Open the calculator