Quantization explained: Q4 vs Q8 and what you lose
Short answer
Quantization stores a model's weights with fewer bits so it uses less memory and runs faster. 8-bit is nearly indistinguishable from the full model, 4-bit (Q4_K_M or MLX 4-bit) is the usual sweet spot with a small quality loss, and below 4-bit quality drops noticeably. A bigger model at 4-bit usually beats a smaller model at 8-bit in the same memory.
Language models are made of billions of numbers called weights. Full-size models store each weight in 16 bits. Quantization rounds those weights to fewer bits, for example 8 or 4, which shrinks the file and the memory it needs.
Why it matters on a Mac
Memory is the limit on what you can run, and speed is limited by how many bytes are read per word generated. A 4-bit model takes about a quarter of the memory of the 16-bit original and runs up to several times faster.
Common formats
| Name you will see | Bits per weight (approx.) | Size vs full model | Quality |
|---|---|---|---|
| F16 / BF16 | 16 | 100% | Original |
| Q8_0, MLX 8-bit | 8.5 | ~53% | Practically identical |
| Q6_K | 6.6 | ~41% | Very close |
| Q5_K_M | 5.7 | ~36% | Very close |
| Q4_K_M, MLX 4-bit | 4.5 to 5 | ~30% | Small loss, the usual default |
| Q3_K_M | 3.9 | ~24% | Noticeable loss |
| Q2_K, IQ2 | 2.5 to 3 | ~17% | Large loss, last resort |
GGUF names (Q4_K_M and so on) are used by llama.cpp, Ollama and LM Studio. MLX models are usually labelled simply 4-bit, 6-bit or 8-bit.
Bigger model, fewer bits
With a fixed amount of memory, you usually get better answers from a larger model at 4-bit than a smaller one at 8-bit. A 14B model at 4-bit (about 8.5 GB) generally outperforms an 8B model at 8-bit (about 8.5 GB). This stops being true at very low bit counts, where quality falls off quickly.
What to pick
- Default: Q4_K_M or MLX 4-bit.
- Spare memory: go up to Q6_K or 8-bit for a small quality gain.
- Coding and math: these suffer most from heavy quantization, so prefer 5-bit or higher when you can.
- Tight memory: pick a smaller model before going below 4-bit.
For the memory each choice needs, see how much memory you need to run an LLM.
Frequently asked questions
What does Q4_K_M mean?
It is a llama.cpp (GGUF) quantization type. Q4 means about 4 bits per weight, K means the k-quant method that groups weights in blocks with shared scales, and M means the medium variant, which keeps some sensitive layers at higher precision.
Which quantization should I download?
Start with Q4_K_M for GGUF or 4-bit for MLX. If you have memory to spare, try Q5_K_M, Q6_K or 8-bit. Only drop to 3-bit or lower when it is the only way to fit a much larger model.