Can Qwen2.5-72B-Instruct run on GeForce RTX 4090 24GB (desktop)?
Conditional memory judgement for one configuration. Change context, quant, or engine and the conclusion can change.
72B does not fit a 24GB card even in Q4 without offload.
Configuration used on this page: gguf-q4_k_m, 4,096 cached tokens, 1 sequence, engine llama.cpp.
Status: Insufficient memory. Software: Documented support.
Weights 40.99 GiB · raw KV 1.25 GiB · total 42.36 GiB–43.58 GiB.
Available budget 22.00 GiB after the default reserve.
This page is an estimate unless a matching runtime-resident measurement is attached.
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-72b-instruct-q4_k_m-00001-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00002-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00003-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00004-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00005-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00006-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00007-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00008-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00009-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00010-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00011-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00012-of-00012.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 8 KV heads, not 64 query heads.
Related: Qwen2.5-72B-Instruct· GeForce RTX 4090 24GB (desktop)
Estimate
Insufficient memory
Even the low estimate exceeds the available budget for this configuration.
Estimated memory range: 42.36 GiB – 43.58 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 22.00 GiB · Remaining after estimate: -21.58 GiB – -20.36 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-72b-instruct-q4_k_m-00001-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00002-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00003-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00004-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00005-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00006-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00007-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00008-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00009-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00010-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00011-of-00012.gguf, qwen2.5-72b-instruct-q4_k_m-00012-of-00012.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 8 KV heads, not 64 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Host RAM is not added to discrete GPU VRAM.