NVIDIA A100 80GB SXM

NVIDIA · A100 80GB SXM HBM2e · server-gpu

Marketing label: 80GB HBM2e.

Capacity used for estimates: 80.00 GiB (85.90 GB).

Default system/display reserve (adjustable): 2.00 GiB.

Vendor-marketed GB for discrete GPU VRAM is treated as GiB (2^30 bytes), matching physical GDDR/HBM module organization. This is an explicit unit assumption.

Memory kind: discrete-vram. Software ecosystem notes: CUDA, vllm.

Catalog models at 8,192 cached tokens

Using each model’s GGUF Q4_K_M file when present, otherwise BF16 files. Concurrent sequences = 1. This is a memory screen, not a performance ranking.

ModelQuantStatusEstimate
Qwen2.5-7B-InstructGGUF Q4_K_Mlikely-fits4.92 GiB–5.42 GiB
Qwen2.5-0.5B-InstructBF16 safetensorslikely-fits1.64 GiB–3.52 GiB
Qwen2.5-1.5B-InstructBF16 safetensorslikely-fits3.72 GiB–5.60 GiB
Qwen2.5-3B-InstructGGUF Q4_K_Mlikely-fits2.37 GiB–2.80 GiB
Qwen2.5-14B-InstructGGUF Q4_K_Mlikely-fits10.00 GiB–10.64 GiB
Qwen2.5-32B-InstructGGUF Q4_K_Mlikely-fits20.61 GiB–21.50 GiB
Qwen2.5-72B-InstructGGUF Q4_K_Mlikely-fits43.61 GiB–44.98 GiB
Qwen3-1.7BBF16 safetensorslikely-fits5.28 GiB–7.19 GiB
Qwen3-4BGGUF Q4_K_Mlikely-fits3.58 GiB–4.08 GiB
Qwen3-8BGGUF Q4_K_Mlikely-fits5.93 GiB–6.48 GiB
Qwen3-14BGGUF Q4_K_Mlikely-fits9.76 GiB–10.39 GiB
Qwen3-32BGGUF Q4_K_Mlikely-fits20.53 GiB–21.41 GiB
Qwen3-30B-A3BGGUF Q4_K_Mlikely-fits18.16 GiB–18.93 GiB
Phi-3.5-mini-instructBF16 safetensorslikely-fits10.74 GiB–12.71 GiB
Mistral-7B-Instruct-v0.3BF16 safetensorslikely-fits15.13 GiB–17.03 GiB
DeepSeek-R1-Distill-Llama-8BBF16 safetensorslikely-fits16.58 GiB–18.49 GiB
DeepSeek-R1-Distill-Qwen-32BBF16 safetensorslikely-fits63.65 GiB–65.59 GiB

Sources

Calculator

Pick a device and workload, then see which catalog models may fit.

Advanced settings

GiB (2^30 bytes)

Custom model structure

Parsed in the browser. The file is not uploaded.

Only huggingface.co, config.json, no token. Private repos are rejected.

Estimate

Likely fits

Under the current assumptions, the high estimate stays below 90% of the available budget.

Estimated memory range: 4.92 GiB5.42 GiB

Weights use selected file metadata. Runtime extras are still estimated.

Weights4.36 GiB
KV cache (raw)0.44 GiB
KV extras (quant metadata / packing)0.00 GiB
Activation / workspace0.06 GiB
Engine reserve0.06 GiB
Safety margin0.00 GiB

Available budget: 78.00 GiB · Remaining after estimate: 72.58 GiB73.08 GiB

Software compatibility

Documented support. Enough memory does not mean this engine, quant, and OS will run. Support does not mean it will be fast.

Assumptions that affect this result

  • Full-resident weight estimates use total parameters, including inactive MoE experts.
  • Weight bytes come from the selected weight files only, not from every file in the repository.
  • Download file size is not peak GPU memory. Runtime layout, allocator padding, and KV are extra.
  • Counted files: qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf, qwen2.5-7b-instruct-q4_k_m-00002-of-00002.gguf.
  • Context budget is the number of cached tokens per sequence (prompt plus reserved generation).
  • Concurrent sequences share one model replica; weights are not multiplied by sequence count.
  • KV dtype is independent of weight quantization unless you change it.
  • Using 4 KV heads, not 28 query heads.
  • Prefill peak and steady decode occupancy are not the same. The high scenario is closer to prefill/workspace pressure.
  • Engine reserve is not the same as bytes the model actually uses for weights and KV.
  • llama.cpp compute buffers grow with context and batch. Values here are user-adjustable scenarios, not measured traces — unless a matching calibration is shown.
  • Host RAM is not added to discrete GPU VRAM.

Calculation method

Catalog models under this workload

ModelEstimateEstimated memory range
Likely fits4.92 GiB5.42 GiB
Likely fits1.14 GiB1.54 GiB
Likely fits3.22 GiB3.67 GiB
Likely fits2.37 GiB2.80 GiB
Likely fits10.00 GiB10.64 GiB
Likely fits20.61 GiB21.50 GiB
Likely fits43.61 GiB44.98 GiB
Likely fits4.78 GiB5.30 GiB
Likely fits3.58 GiB4.08 GiB
Likely fits5.93 GiB6.48 GiB
Likely fits9.76 GiB10.39 GiB
Likely fits20.53 GiB21.41 GiB
Likely fits18.16 GiB18.93 GiB
Likely fits10.24 GiB10.97 GiB
Likely fits14.63 GiB15.34 GiB
Likely fits16.08 GiB16.83 GiB
Likely fits63.15 GiB64.89 GiB