NVIDIA DGX Spark (GB10, 128 GB unified)

NVIDIA · NVIDIA DGX Spark 940-54242-0000 · unified-device

Marketing label: 128 GB LPDDR5x coherent unified system memory.

Capacity used for estimates: 128.00 GiB (137.44 GB).

Default system/display reserve (adjustable): 16.00 GiB.

Marketed 128 GB unified system memory is treated as 128 GiB (2^30 bytes) for calculation. Users can override the budget. Unified memory is a shared CPU/GPU pool; host RAM is not added again.

Memory kind: unified-memory. Software ecosystem notes: CUDA, vllm, NVIDIA AI software stack.

Catalog models at 8,192 cached tokens

Using each model’s GGUF Q4_K_M file when present, otherwise BF16 files. Concurrent sequences = 1. This is a memory screen, not a performance ranking.

ModelQuantStatusEstimate
Qwen2.5-7B-InstructGGUF Q4_K_Mlikely-fits4.92 GiB–5.42 GiB
Qwen2.5-0.5B-InstructBF16 safetensorslikely-fits1.64 GiB–3.52 GiB
Qwen2.5-1.5B-InstructBF16 safetensorslikely-fits3.72 GiB–5.60 GiB
Qwen2.5-3B-InstructGGUF Q4_K_Mlikely-fits2.37 GiB–2.80 GiB
Qwen2.5-14B-InstructGGUF Q4_K_Mlikely-fits10.00 GiB–10.64 GiB
Qwen2.5-32B-InstructGGUF Q4_K_Mlikely-fits20.61 GiB–21.50 GiB
Qwen2.5-72B-InstructGGUF Q4_K_Mlikely-fits43.61 GiB–44.98 GiB
Qwen3-1.7BBF16 safetensorslikely-fits5.28 GiB–7.19 GiB
Qwen3-4BGGUF Q4_K_Mlikely-fits3.58 GiB–4.08 GiB
Qwen3-8BGGUF Q4_K_Mlikely-fits5.93 GiB–6.48 GiB
Qwen3-14BGGUF Q4_K_Mlikely-fits9.76 GiB–10.39 GiB
Qwen3-32BGGUF Q4_K_Mlikely-fits20.53 GiB–21.41 GiB
Qwen3-30B-A3BGGUF Q4_K_Mlikely-fits18.16 GiB–18.93 GiB
Phi-3.5-mini-instructBF16 safetensorslikely-fits10.74 GiB–12.71 GiB
Mistral-7B-Instruct-v0.3BF16 safetensorslikely-fits15.13 GiB–17.03 GiB
Mixtral-8x7B-Instruct-v0.1BF16 safetensorslikely-fits88.62 GiB–90.52 GiB
DeepSeek-R1-Distill-Llama-8BBF16 safetensorslikely-fits16.58 GiB–18.49 GiB
DeepSeek-R1-Distill-Qwen-32BBF16 safetensorslikely-fits63.65 GiB–65.59 GiB

Sources

Calculator

Pick a device and workload, then see which catalog models may fit.

Advanced settings

GiB (2^30 bytes)

Custom model structure

Parsed in the browser. The file is not uploaded.

Only huggingface.co, config.json, no token. Private repos are rejected.

Estimate

Likely fits

Under the current assumptions, the high estimate stays below 90% of the available budget.

Estimated memory range: 4.92 GiB5.42 GiB

Weights use selected file metadata. Runtime extras are still estimated.

Weights4.36 GiB
KV cache (raw)0.44 GiB
KV extras (quant metadata / packing)0.00 GiB
Activation / workspace0.06 GiB
Engine reserve0.06 GiB
Safety margin0.00 GiB

Available budget: 112.00 GiB · Remaining after estimate: 106.58 GiB107.08 GiB

Software compatibility

Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.

Assumptions that affect this result

  • Full-resident weight estimates use total parameters, including inactive MoE experts.
  • Weight bytes come from the selected weight files only, not from every file in the repository.
  • Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
  • Counted files: qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf, qwen2.5-7b-instruct-q4_k_m-00002-of-00002.gguf.
  • Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
  • Concurrent sequences share one model replica; weights are not multiplied by sequence count.
  • KV dtype is independent of weight quantization unless you change it.
  • Using 4 KV heads, not 28 query heads.
  • Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
  • Engine reserve is not the same as bytes the model actually uses for weights and KV.
  • llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
  • Unified memory is a shared CPU/GPU pool. Host RAM is not added on top of this capacity.

Calculation method

Catalog models under this workload

ModelEstimateEstimated memory range
Likely fits4.92 GiB5.42 GiB
Likely fits1.14 GiB1.54 GiB
Likely fits3.22 GiB3.67 GiB
Likely fits2.37 GiB2.80 GiB
Likely fits10.00 GiB10.64 GiB
Likely fits20.61 GiB21.50 GiB
Likely fits43.61 GiB44.98 GiB
Likely fits4.78 GiB5.30 GiB
Likely fits3.58 GiB4.08 GiB
Likely fits5.93 GiB6.48 GiB
Likely fits9.76 GiB10.39 GiB
Likely fits20.53 GiB21.41 GiB
Likely fits18.16 GiB18.93 GiB
Likely fits10.24 GiB10.97 GiB
Likely fits14.63 GiB15.34 GiB
Likely fits88.12 GiB90.30 GiB
Likely fits16.08 GiB16.83 GiB
Likely fits63.15 GiB64.89 GiB