Can Qwen2.5-3B-Instruct run on GeForce RTX 4090 24GB (desktop)?
Conditional memory judgement for one configuration. Change context, quant, or engine and the conclusion can change.
3B Q4_K_M on RTX 4090 with a matching llama.cpp runtime-resident trace at 8192 tokens.
Configuration used on this page: gguf-q4_k_m, 8,192 cached tokens, 1 sequence, engine llama.cpp.
Status: Likely fits. Software: Documented support.
Weights 1.96 GiB · raw KV 0.28 GiB · total 2.37 GiB–2.80 GiB.
Available budget 22.00 GiB after the default reserve.
Runtime-resident measurement rt-qwen2.5-3b-instruct-q4km-4090-ctx8192: process GPU 2.59 GiB on 2026-09-19T15:43:43Z (llama.cpp version: 0.4.1-dev (build 1, commit 1af554f)). Formula estimate remains shown separately.
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-3b-instruct-q4_k_m.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 2 KV heads, not 16 query heads.
Related: Qwen2.5-3B-Instruct· GeForce RTX 4090 24GB (desktop)
Estimate
Likely fits
Under the current assumptions, the high estimate stays below 90% of the available budget.
Estimated memory range: 2.37 GiB – 2.80 GiB
Runtime-resident measurement for a matching environment.
Available budget: 22.00 GiB · Remaining after estimate: 19.20 GiB – 19.63 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Runtime-resident measurement for a matching environment. Measured process GPU memory: 2.59 GiB · llama.cpp version: 0.4.1-dev (build 1, commit 1af554f) · 2026-09-19T15:43:43Z
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-3b-instruct-q4_k_m.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 2 KV heads, not 16 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Host RAM is not added to discrete GPU VRAM.