Can DeepSeek-V2-Lite-Chat run on GeForce RTX 4090 24GB (desktop)?
Conditional memory judgement for one configuration. Change context, quant, or engine and the conclusion can change.
MLA is unsupported for full KV; page explains the gap instead of faking GQA.
Configuration used on this page: bf16, 4,096 cached tokens, 1 sequence, engine vllm.
Status: Unsupported / incomplete. Software: Documented support.
Weights 29.26 GiB · raw KV 0.00 GiB · total 29.26 GiB–29.26 GiB.
Available budget 22.00 GiB after the default reserve.
This page is an estimate unless a matching runtime-resident measurement is attached.
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- This MoE lists 15700000000 total parameters and 2400000000 active parameters. Active-parameter compute is not used as the weight-residency figure.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: model.safetensors.index.json metadata.total_size.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
Related: DeepSeek-V2-Lite-Chat· GeForce RTX 4090 24GB (desktop)
Estimate
Unsupported / incomplete
The architecture or split mode is out of scope. Weight figures may still be shown.
Estimated memory range: 29.26 GiB – 29.26 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 22.00 GiB · Remaining after estimate: -7.26 GiB – -7.26 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
- Architecture mla is not fully supported for KV estimates.
- MLA KV is not estimated with the GQA formula.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- This MoE lists 15700000000 total parameters and 2400000000 active parameters. Active-parameter compute is not used as the weight-residency figure.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: model.safetensors.index.json metadata.total_size.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Host RAM is not added to discrete GPU VRAM.