GeForce RTX 4060 8GB (desktop)
NVIDIA · GeForce RTX 4060 8GB desktop · desktop-gpu
Marketing label: 8 GB GDDR6.
Capacity used for estimates: 8.00 GiB (8.59 GB).
Default system/display reserve (adjustable): 1.00 GiB.
Vendor-marketed GB for discrete GPU VRAM is treated as GiB (2^30 bytes), matching physical GDDR/HBM module organization. This is an explicit unit assumption.
Memory kind: discrete-vram. Software ecosystem notes: CUDA, llama.cpp, ollama, vllm.
Catalog models at 8,192 cached tokens
Using each model’s GGUF Q4_K_M file when present, otherwise BF16 files. Concurrent sequences = 1. This is a memory screen, not a performance ranking.
| Model | Quant | Status | Estimate |
|---|---|---|---|
| Qwen2.5-7B-Instruct | GGUF Q4_K_M | likely-fits | 4.92 GiB–5.42 GiB |
| Qwen2.5-0.5B-Instruct | BF16 safetensors | likely-fits | 1.64 GiB–3.52 GiB |
| Qwen2.5-1.5B-Instruct | BF16 safetensors | likely-fits | 3.72 GiB–5.60 GiB |
| Qwen2.5-3B-Instruct | GGUF Q4_K_M | likely-fits | 2.37 GiB–2.80 GiB |
| Qwen3-1.7B | BF16 safetensors | tight | 5.28 GiB–7.19 GiB |
| Qwen3-4B | GGUF Q4_K_M | likely-fits | 3.58 GiB–4.08 GiB |
| Qwen3-8B | GGUF Q4_K_M | tight | 5.93 GiB–6.48 GiB |
Sources
- NVIDIA GeForce RTX 40 Series specs — Memory Size 8 GB GDDR6
Calculator
Estimate
Likely fits
Under the current assumptions, the high estimate stays below 90% of the available budget.
Estimated memory range: 4.92 GiB – 5.42 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 7.00 GiB · Remaining after estimate: 1.58 GiB – 2.08 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf, qwen2.5-7b-instruct-q4_k_m-00002-of-00002.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 4 KV heads, not 28 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Host RAM is not added to discrete GPU VRAM.
Catalog models under this workload
| Model | Estimate | Estimated memory range |
|---|---|---|
| Likely fits | 4.92 GiB – 5.42 GiB | |
| Likely fits | 1.14 GiB – 1.54 GiB | |
| Likely fits | 3.22 GiB – 3.67 GiB | |
| Likely fits | 2.37 GiB – 2.80 GiB | |
| Likely fits | 4.78 GiB – 5.30 GiB | |
| Likely fits | 3.58 GiB – 4.08 GiB | |
| Tight | 5.93 GiB – 6.48 GiB |