GGUF quant families

A GGUF name encodes which generation of algorithm produced it. Legacy quants end in _0 or _1 and give each block of 32 weights its own fp16 constants. K-quants add a _K, group eight blocks into a 256-weight super-block, and quantize the block constants themselves so the bookkeeping shrinks. I-quants use IQ and abandon buckets entirely, matching groups of eight weights against a fixed codebook of reference vectors. S, M, and L are not : they say how many tensors were hand-flagged for extra precision, which is why a Q4_K_M file measures above its 4.5 bpw block rate. An is orthogonal to all three and costs nothing in size.

Formula

Worked example: Taking a Q4_K super-block apart

256 weights at 4 bits is 128 bytes. Q4_K adds 12 bytes of packed 6-bit scales and mins, then 4 bytes of fp16 constants that dequantize those, for 144 total. The extra 16 bytes are exactly the half bit that separates a nominal Q4 from the 4.5 bpw this simulator charges. Legacy Q4_0 spends its constants per 32 weights instead of per 256, which is the whole difference between the generations.

Try Q4_K_M and read the memory breakdown