NVIDIA lists GeForce RTX 4070 as 12 GB, RTX 4080 SUPER as 16 GB, and RTX 4090 as 24 GB GDDR6X. This catalog stores those as 12, 16, and 24 GiB of discrete VRAM, because GDDR capacities are binary modules. The assumption is written on each hardware page. Host RAM is not added to those numbers.
Default reserves in this dataset: 1 GiB on 12–16 GB cards, 2 GiB on the 24 GB 4090. If you already typed a net budget, the reserve is not subtracted again.
Worked screens, calculator 0.1.0, 8,192 tokens, 1 sequence, FP16 KV, llama.cpp typical extras unless noted:
- Qwen2.5-7B-Instruct Q4_K_M on an RTX 4060 8 GB (7 GiB net after 1 GiB reserve): weights ~4.36 GiB plus small KV. Often “likely fits” or “tight” depending on extras — not a guarantee the OS, display, and browser leave that reserve free.
- Qwen2.5-14B-Instruct Q4_K_M (8,988,110,496 bytes of those three shards) on a 16 GB 4060 Ti SKU: weights ~8.37 GiB before KV.
- Qwen2.5-32B-Instruct Q4_K_M (19,851,336,384 bytes) versus a 24 GB 4090 (22 GiB net): the weight file alone is ~18.5 GiB, so long context makes the status tight or insufficient.
Always compare the range to the net budget. Comparing only the GGUF size to “24GB” on the box is the usual way to get a surprise OOM.