NVIDIA DGX Spark (GB10, 128 GB unified)
NVIDIA · NVIDIA DGX Spark 940-54242-0000 · unified-device
Marketing label: 128 GB LPDDR5x coherent unified system memory.
Capacity used for estimates: 128.00 GiB (137.44 GB).
Default system/display reserve (adjustable): 16.00 GiB.
Marketed 128 GB unified system memory is treated as 128 GiB (2^30 bytes) for calculation. Users can override the budget. Unified memory is a shared CPU/GPU pool; host RAM is not added again.
Memory kind: unified-memory. Software ecosystem notes: CUDA, vllm, NVIDIA AI software stack.
Catalog models at 8,192 cached tokens
Using each model’s GGUF Q4_K_M file when present, otherwise BF16 files. Concurrent sequences = 1. This is a memory screen, not a performance ranking.
| Model | Quant | Status | Estimate |
|---|---|---|---|
| Qwen2.5-7B-Instruct | GGUF Q4_K_M | likely-fits | 4.92 GiB–5.42 GiB |
| Qwen2.5-0.5B-Instruct | BF16 safetensors | likely-fits | 1.64 GiB–3.52 GiB |
| Qwen2.5-1.5B-Instruct | BF16 safetensors | likely-fits | 3.72 GiB–5.60 GiB |
| Qwen2.5-3B-Instruct | GGUF Q4_K_M | likely-fits | 2.37 GiB–2.80 GiB |
| Qwen2.5-14B-Instruct | GGUF Q4_K_M | likely-fits | 10.00 GiB–10.64 GiB |
| Qwen2.5-32B-Instruct | GGUF Q4_K_M | likely-fits | 20.61 GiB–21.50 GiB |
| Qwen2.5-72B-Instruct | GGUF Q4_K_M | likely-fits | 43.61 GiB–44.98 GiB |
| Qwen3-1.7B | BF16 safetensors | likely-fits | 5.28 GiB–7.19 GiB |
| Qwen3-4B | GGUF Q4_K_M | likely-fits | 3.58 GiB–4.08 GiB |
| Qwen3-8B | GGUF Q4_K_M | likely-fits | 5.93 GiB–6.48 GiB |
| Qwen3-14B | GGUF Q4_K_M | likely-fits | 9.76 GiB–10.39 GiB |
| Qwen3-32B | GGUF Q4_K_M | likely-fits | 20.53 GiB–21.41 GiB |
| Qwen3-30B-A3B | GGUF Q4_K_M | likely-fits | 18.16 GiB–18.93 GiB |
| Phi-3.5-mini-instruct | BF16 safetensors | likely-fits | 10.74 GiB–12.71 GiB |
| Mistral-7B-Instruct-v0.3 | BF16 safetensors | likely-fits | 15.13 GiB–17.03 GiB |
| Mixtral-8x7B-Instruct-v0.1 | BF16 safetensors | likely-fits | 88.62 GiB–90.52 GiB |
| DeepSeek-R1-Distill-Llama-8B | BF16 safetensors | likely-fits | 16.58 GiB–18.49 GiB |
| DeepSeek-R1-Distill-Qwen-32B | BF16 safetensors | likely-fits | 63.65 GiB–65.59 GiB |
Sources
- NVIDIA DGX Spark product page — 128 GB LPDDR5x coherent unified system memory. This is not dedicated VRAM. Default reserve deducts OS and other processes. Multiple Sparks are not one pool.
Calculator
Estimate
Likely fits
Under the current assumptions, the high estimate stays below 90% of the available budget.
Estimated memory range: 4.92 GiB – 5.42 GiB
Weights use selected file metadata. Runtime extras are still estimated.
Available budget: 112.00 GiB · Remaining after estimate: 106.58 GiB – 107.08 GiB
Software compatibility
Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.
Assumptions that affect this result
- Full-resident weight estimates use total parameters, including inactive MoE experts.
- Weight bytes come from the selected weight files only, not from every file in the repository.
- Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
- Counted files: qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf, qwen2.5-7b-instruct-q4_k_m-00002-of-00002.gguf.
- Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
- Concurrent sequences share one model replica; weights are not multiplied by sequence count.
- KV dtype is independent of weight quantization unless you change it.
- Using 4 KV heads, not 28 query heads.
- Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
- Engine reserve is not the same as bytes the model actually uses for weights and KV.
- llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
- Unified memory is a shared CPU/GPU pool. Host RAM is not added on top of this capacity.
Catalog models under this workload
| Model | Estimate | Estimated memory range |
|---|---|---|
| Likely fits | 4.92 GiB – 5.42 GiB | |
| Likely fits | 1.14 GiB – 1.54 GiB | |
| Likely fits | 3.22 GiB – 3.67 GiB | |
| Likely fits | 2.37 GiB – 2.80 GiB | |
| Likely fits | 10.00 GiB – 10.64 GiB | |
| Likely fits | 20.61 GiB – 21.50 GiB | |
| Likely fits | 43.61 GiB – 44.98 GiB | |
| Likely fits | 4.78 GiB – 5.30 GiB | |
| Likely fits | 3.58 GiB – 4.08 GiB | |
| Likely fits | 5.93 GiB – 6.48 GiB | |
| Likely fits | 9.76 GiB – 10.39 GiB | |
| Likely fits | 20.53 GiB – 21.41 GiB | |
| Likely fits | 18.16 GiB – 18.93 GiB | |
| Likely fits | 10.24 GiB – 10.97 GiB | |
| Likely fits | 14.63 GiB – 15.34 GiB | |
| Likely fits | 88.12 GiB – 90.30 GiB | |
| Likely fits | 16.08 GiB – 16.83 GiB | |
| Likely fits | 63.15 GiB – 64.89 GiB |