Can Qwen3-30B-A3B run on GeForce RTX 4090 24GB (desktop)?
Conditional memory judgement for one configuration. Change context, quant, or engine and the conclusion can change.
MoE full-resident weights vs 24GB, not active-parameter compute.
Configuration used on this page: gguf-q4_k_m, 8,192 cached tokens, 1 sequence, engine llama.cpp.
Status: Likely fits. Software: Documented support.
Weights 17.28 GiB · raw KV 0.75 GiB · total 18.16 GiB–18.93 GiB.
Available budget 22.00 GiB after the default reserve.
This page is an estimate unless a matching runtime-resident measurement is attached.
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- This MoE lists 30500000000 total parameters and 3300000000 active parameters. Active-parameter compute is not used as the weight-residency figure.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: Qwen3-30B-A3B-Q4_K_M.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
Related: Qwen3-30B-A3B· GeForce RTX 4090 24GB (desktop)
Estimate
Likely fits
Under the current assumptions, the high estimate stays below 90% of the available budget.
Estimated memory range: 18.16 GiB – 18.93 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 22.00 GiB · Remaining after estimate: 3.07 GiB – 3.84 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- This MoE lists 30500000000 total parameters and 3300000000 active parameters. Active-parameter compute is not used as the weight-residency figure.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: Qwen3-30B-A3B-Q4_K_M.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 4 KV heads, not 32 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Host RAM is not added to discrete GPU VRAM.