Can Qwen2.5-3B-Instruct run on GeForce RTX 4090 24GB (desktop)?

Conditional memory judgement for one configuration. Change context, quant, or engine and the conclusion can change.

3B Q4_K_M on RTX 4090 with a matching llama.cpp runtime-resident trace at 8192 tokens.

Configuration used on this page: gguf-q4_k_m, 8,192 cached tokens, 1 sequence, engine llama.cpp.

Status: Likely fits. Software: Documented support.

Weights 1.96 GiB · raw KV 0.28 GiB · total 2.37 GiB–2.80 GiB.

Available budget 22.00 GiB after the default reserve.

Runtime-resident measurement rt-qwen2.5-3b-instruct-q4km-4090-ctx8192: process GPU 2.59 GiB on 2026-09-19T15:43:43Z (llama.cpp version: 0.4.1-dev (build 1, commit 1af554f)). Formula estimate remains shown separately.

Related: Qwen2.5-3B-Instruct· GeForce RTX 4090 24GB (desktop)

Pick one model and one device for a conditional judgement.

Advanced settings

GiB (2^30 bytes)

Custom model structure

Parsed in the browser. The file is not uploaded.

Only huggingface.co, config.json, no token. Private repos are rejected.

Estimate

Likely fits

Under the current assumptions, the high estimate stays below 90% of the available budget.

Estimated memory range: 2.37 GiB2.80 GiB

Runtime-resident measurement for a matching environment.

Weights1.96 GiB
KV cache (raw)0.28 GiB
KV extras (quant metadata / packing)0.00 GiB
Activation / workspace0.06 GiB
Engine reserve0.06 GiB
Safety margin0.00 GiB

Available budget: 22.00 GiB · Remaining after estimate: 19.20 GiB19.63 GiB

Software compatibility

Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.

Runtime-resident measurement for a matching environment. Measured process GPU memory: 2.59 GiB · llama.cpp version: 0.4.1-dev (build 1, commit 1af554f) · 2026-09-19T15:43:43Z

Assumptions that affect this result

  • Full-resident weight estimates use total parameters, including inactive MoE experts.
  • Weight bytes come from the selected weight files only, not from every file in the repository.
  • Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
  • Counted files: qwen2.5-3b-instruct-q4_k_m.gguf.
  • Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
  • Concurrent sequences share one model replica; weights are not multiplied by sequence count.
  • KV dtype is independent of weight quantization unless you change it.
  • Using 2 KV heads, not 16 query heads.
  • Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
  • Engine reserve is not the same as bytes the model actually uses for weights and KV.
  • llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
  • Host RAM is not added to discrete GPU VRAM.

Calculation method