Models

Specific checkpoints with sourced architecture and weight files.

ModelParametersAttentionSupportQ4_K_M or BF16 @ 8k
Qwen2.5-7B-Instruct7.61Bgqafull4.92 GiB
Qwen2.5-0.5B-Instruct0.49Bgqafull1.08 GiB
Qwen2.5-1.5B-Instruct1.54Bgqafull3.16 GiB
Qwen2.5-3B-Instruct3.09Bgqafull2.37 GiB
Qwen2.5-14B-Instruct14.70Bgqafull10.00 GiB
Qwen2.5-32B-Instruct32.50Bgqafull20.61 GiB
Qwen2.5-72B-Instruct72.70Bgqafull43.61 GiB
Qwen3-1.7B1.70Bgqafull4.72 GiB
Qwen3-4B4.00Bgqafull3.58 GiB
Qwen3-8B8.20Bgqafull5.93 GiB
Qwen3-14B14.80Bgqafull9.76 GiB
Qwen3-32B32.80Bgqafull20.53 GiB
Qwen3-30B-A3B30.50Bgqafull18.16 GiB
Phi-3.5-mini-instruct3.82Bmhafull10.18 GiB
Mistral-7B-Instruct-v0.37.25Bgqafull14.56 GiB
Mixtral-8x7B-Instruct-v0.146.70Bgqafull88.05 GiB
DeepSeek-R1-Distill-Llama-8B8.03Bgqafull16.02 GiB
DeepSeek-R1-Distill-Qwen-32B32.50Bgqafull63.09 GiB
DeepSeek-V2-Lite-Chat15.70Bmlaunsupported29.26 GiB