Quantization formats

A quantization format is more than a bit count. It also describes how weights are grouped, scaled, and sometimes protected according to their importance. Q8_0 is near-lossless, Q4_K_M is the practical llama.cpp default, and activation-aware formats such as AWQ target . Extremely low-bit formats need a large model or weights trained natively at that precision.

Formula

artifact size (GB) = exact file bytes ÷ 1,000,000,000
raw weight estimate (GB) = parameters (billions) × bpw ÷ 8

Worked example: Qwen3-32B in Q4_K_M

The research artifact records an exact bartowski Q4_K_M file size of 19,762,149,696 bytes. It also places Q4_K_M at about 0.8% perplexity change from the full-precision baseline, while warning that quality loss becomes visible below Q3.

19,762,149,696 bytes ÷ 1,000,000,000 = 19.76 GB
Try Q4_K_M on a current 27B model