GeForce RTX 4090 24GB (desktop)
NVIDIA · GeForce RTX 4090 Founders Edition / 24GB desktop · desktop-gpu
Marketing label: 24 GB GDDR6X.
Capacity used for estimates: 24.00 GiB (25.77 GB).
Default system/display reserve (adjustable): 2.00 GiB.
Vendor-marketed GB for discrete GPU VRAM is treated as GiB (2^30 bytes), matching physical GDDR/HBM module organization. This is an explicit unit assumption.
Memory kind: discrete-vram. Software ecosystem notes: CUDA, llama.cpp, ollama, vllm.
Catalog models at 8,192 cached tokens
Using each model’s GGUF Q4_K_M file when present, otherwise BF16 files. Concurrent sequences = 1. This is a memory screen, not a performance ranking.
| Model | Quant | Status | Estimate |
|---|---|---|---|
| Qwen2.5-7B-Instruct | GGUF Q4_K_M | likely-fits | 4.92 GiB–5.42 GiB |
| Qwen2.5-0.5B-Instruct | BF16 safetensors | likely-fits | 1.64 GiB–3.52 GiB |
| Qwen2.5-1.5B-Instruct | BF16 safetensors | likely-fits | 3.72 GiB–5.60 GiB |
| Qwen2.5-3B-Instruct | GGUF Q4_K_M | likely-fits | 2.37 GiB–2.80 GiB |
| Qwen2.5-14B-Instruct | GGUF Q4_K_M | likely-fits | 10.00 GiB–10.64 GiB |
| Qwen2.5-32B-Instruct | GGUF Q4_K_M | tight | 20.61 GiB–21.50 GiB |
| Qwen3-1.7B | BF16 safetensors | likely-fits | 5.28 GiB–7.19 GiB |
| Qwen3-4B | GGUF Q4_K_M | likely-fits | 3.58 GiB–4.08 GiB |
| Qwen3-8B | GGUF Q4_K_M | likely-fits | 5.93 GiB–6.48 GiB |
| Qwen3-14B | GGUF Q4_K_M | likely-fits | 9.76 GiB–10.39 GiB |
| Qwen3-32B | GGUF Q4_K_M | tight | 20.53 GiB–21.41 GiB |
| Qwen3-30B-A3B | GGUF Q4_K_M | likely-fits | 18.16 GiB–18.93 GiB |
| Phi-3.5-mini-instruct | BF16 safetensors | likely-fits | 10.74 GiB–12.71 GiB |
| Mistral-7B-Instruct-v0.3 | BF16 safetensors | likely-fits | 15.13 GiB–17.03 GiB |
| DeepSeek-R1-Distill-Llama-8B | BF16 safetensors | likely-fits | 16.58 GiB–18.49 GiB |
Sources
- NVIDIA GeForce RTX 4090 product page — Standard Memory Config 24 GB GDDR6X
Calculator
Estimate
Likely fits
Under the current assumptions, the high estimate stays below 90% of the available budget.
Estimated memory range: 4.92 GiB – 5.42 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 22.00 GiB · Remaining after estimate: 16.58 GiB – 17.08 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf, qwen2.5-7b-instruct-q4_k_m-00002-of-00002.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 4 KV heads, not 28 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Host RAM is not added to discrete GPU VRAM.
Catalog models under this workload
| Model | Estimate | Estimated memory range |
|---|---|---|
| Likely fits | 4.92 GiB – 5.42 GiB | |
| Likely fits | 1.14 GiB – 1.54 GiB | |
| Likely fits | 3.22 GiB – 3.67 GiB | |
| Likely fits | 2.37 GiB – 2.80 GiB | |
| Likely fits | 10.00 GiB – 10.64 GiB | |
| Tight | 20.61 GiB – 21.50 GiB | |
| Likely fits | 4.78 GiB – 5.30 GiB | |
| Likely fits | 3.58 GiB – 4.08 GiB | |
| Likely fits | 5.93 GiB – 6.48 GiB | |
| Likely fits | 9.76 GiB – 10.39 GiB | |
| Tight | 20.53 GiB – 21.41 GiB | |
| Likely fits | 18.16 GiB – 18.93 GiB | |
| Likely fits | 10.24 GiB – 10.97 GiB | |
| Likely fits | 14.63 GiB – 15.34 GiB | |
| Likely fits | 16.08 GiB – 16.83 GiB |