Qwen3-1.7B

Qwen checkpoint Qwen/Qwen3-1.7B at revision 70d244cc86cc. Architecture qwen3, attention gqa, calculator support full.

Total parameters: 1,700,000,000.

Layers 28, query heads 16, KV heads 8, head dim 128, recorded context 40,960 tokens.

Fetched 2026-09-19, verified 2026-09-19. Completeness: complete.

Memory by quantization and context

Concurrent sequences = 1. KV cache defaults to FP16. Range includes documented runtime extras for the listed engine. These HTML figures use the same core as the interactive calculator.

Quantization2,048 tokens8,192 tokens32,768 tokens
BF16 safetensors4.63 GiB–6.51 GiB5.28 GiB–7.19 GiB7.91 GiB–9.89 GiB
GGUF Q8_02.05 GiB–2.47 GiB2.71 GiB–3.18 GiB5.33 GiB–5.99 GiB
INT4 theoretical1.14 GiB–1.61 GiB1.79 GiB–2.32 GiB4.42 GiB–5.13 GiB

Sources

Calculator

Pick a model and settings to see memory demand.

Advanced settings

GiB (2^30 bytes)

Custom model structure

Parsed in the browser. The file is not uploaded.

Only huggingface.co, config.json, no token. Private repos are rejected.

Estimate

Unknown

Missing budget, incomplete data, or a deployment mode this version cannot judge.

Estimated memory range: 4.78 GiB5.30 GiB

Weights use selected file metadata. Runtime extras are still estimated.

Weights3.78 GiB
KV cache (raw)0.88 GiB
KV extras (quant metadata / packing)0.00 GiB
Activation / workspace0.06 GiB
Engine reserve0.06 GiB
Safety margin0.00 GiB

Software compatibility

Not verified. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.

  • Memory headroom does not prove this engine, quantization, and OS combination is supported.

Assumptions that affect this result

  • Full-resident weight estimates use total parameters, including inactive MoE experts.
  • Weight bytes come from the selected weight files only, not from every file in the repository.
  • Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
  • Counted files: model.safetensors.index.json metadata.total_size.
  • Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
  • Concurrent sequences share one model replica; weights are not multiplied by sequence count.
  • KV dtype is independent of weight quantization unless you change it.
  • Using 8 KV heads, not 16 query heads.
  • Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
  • Engine reserve is not the same as bytes the model actually uses for weights and KV.
  • llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.

Calculation method

Try these changes