## Quantisation

*Production and serving*

*Last edited: 28 September 2026*

Quantisation stores the weights, and sometimes the activations and the KV cache too, in fewer bits by rounding them to a coarser grid of values. The model takes 2–4 times less memory, and generation gets faster because decode reads fewer bytes from memory.

**In plain words:** A photo saved as a JPEG instead of at full quality: it is several times smaller and looks the same on a phone screen. Artefacts appear only under heavy compression, and first in the fine detail. For a model, the fine detail is long reasoning, code and less common languages.

*Interactive widget on the page: Lower the bit count, add an outlier, then turn on a separate scale per block. Watch the rounding error.*

*Interactive widget on the page: Pick a model size and see what hardware the weights alone fit on. Change the bit count in the widget above.*

### How quantisation works

- Each weight is divided by a scale, rounded to a number from a small range (−8 to 7 in INT4; a symmetric grid, as in the widget, uses −7 to 7, i.e. 15 levels) and stored like that. At compute time it is multiplied back by the scale. The error is rounding noise, and the network tolerates small noise well.
- Outliers are the problem. One large value stretches the grid, and the rest of the weights land on a few points around zero. That is why the scale is computed separately for small groups, typically 16–128 weights. In the activations of models from about 6.7 billion parameters up, systematic outliers appear in a few channels (Dettmers et al., LLM.int8(), 2022).
- Post-training quantisation (PTQ) works on a finished model in minutes or hours. GPTQ and AWQ use a small sample of calibration data to round the weights so that the layer’s output changes as little as possible. Quantisation-aware training (QAT) produces a model that copes better with 4 bits and below, but it means extra training, so it is usually done by the model’s author.

### Quantisation formats and what they speed up

- Weight-only quantisation, e.g. W4A16 (weights in 4 bits, activations in 16), reduces the bytes read in decode, so it speeds up generation at small batch sizes. The multiplication still runs in 16 bits, so prefill and large batches gain almost nothing.
- FP8 for weights and activations (W8A8) lets the tensor cores compute in FP8: an H100 reaches about 1,979 dense TFLOPS versus 989 in BF16. So it also speeds up prefill and large batches. Kurtic et al. (2024) measure FP8 W8A8 as practically lossless, INT8 W8A8 at 1–3% loss, and W4A16 as closer to 8-bit quality than expected.
- The new 4-bit formats are floating-point numbers with a scale per block. MXFP4 uses blocks of 32 weights with a power-of-two scale, about 4.25 bits per weight; OpenAI released the expert weights of gpt-oss in this format. NVFP4 uses blocks of 16 with an FP8 scale plus an extra per-tensor scale, about 4.5 bits, and has native hardware support on Blackwell GPUs.
- The KV cache gets quantised too. An FP8 cache takes half the memory per token, so with long contexts it frees more room for the batch than squeezing the weights further (see “The generation loop and KV cache”).

### What you lose and how to choose a format

- Rule of thumb, depending on the model and method: 8 bits is practically lossless, 4 bits with a good method is a small loss, and at 3 bits and below quality drops off fast. Losses show up first in long reasoning, code, maths and less common languages, and a benchmark average can hide them.
- For the same memory, a larger model in 4 bits usually beats a smaller one in 16. After more than 35,000 experiments, Dettmers and Zettlemoyer (2022) found 4 bits almost always optimal for a fixed total number of model bits.
- The choice depends on traffic and hardware. On Hopper GPUs such as the H100, Kurtic et al. recommend W4A16 for single requests and small batches, and FP8 W8A8 for heavy traffic with continuous batching. On Blackwell, NVFP4 for weights and activations (W4A4) runs on the FP4 tensor cores at 2–3× the FP8 rate, so 4 bits can win under heavy traffic too, at a somewhat larger loss than FP8: Red Hat measures about 99% of BF16 accuracy for 70B+ models and 95–98% for 7–14B ones. Check quality on your own eval set, because the same model can be quantised differently by different hosts (see “Open-weight models” and “Evals”).

### Check yourself

**Question:** What does quantisation give you, what do you lose, and which format would you choose?

**Short answer:** The weights are stored in fewer bits, from 16 down to 8 or 4, with a scale computed per small block, because a single outlier stretches the grid. The model needs 2–4 times less memory, and decode gets faster because it reads fewer bytes. Weight-only quantisation such as W4A16 helps at small batch sizes. FP8 for weights and activations also speeds up the maths, so on Hopper it wins under heavy traffic; on Blackwell, NVFP4 for weights and activations is faster still. 8 bits is practically lossless, 4 bits costs a little, and below that quality drops fast, first in reasoning and code. Your own evals decide.

### Follow-up questions

- **PTQ or QAT?** PTQ quantises a finished model in minutes or hours, using a small sample of calibration data (GPTQ, AWQ). QAT simulates quantisation during training, so the model learns to tolerate it and holds its quality better at 4 bits and below. The cost is training, so QAT is usually done by the model’s author.
- **Why are activations harder to quantise than weights?** Weights are known in advance and can be carefully rescaled offline. Activations depend on the input and have large outliers in a few channels. Methods that shift the difficulty onto the weights (SmoothQuant) help, as does FP8, which has a wider range than INT8.
- **After quantising to 4 bits the benchmarks look fine, but users complain about code. What do you do?** A benchmark average hides losses in long reasoning chains and code. Compare both versions on your own eval set of coding tasks. If the loss is real, go back to 8 bits or keep the sensitive layers at higher precision.
- **When does quantising the KV cache help more than quantising the weights?** With long contexts and large batches, when the cache of all conversations takes more memory than the weights. An FP8 cache holds twice as many tokens, i.e. twice as many conversations or twice the context length.

### Sources

- [Maarten Grootendorst: A Visual Guide to Quantization](https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-quantization)
- [Kurtic et al.: “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization](https://arxiv.org/abs/2411.02355)
- [Dettmers and Zettlemoyer: The case for 4-bit precision (2022)](https://arxiv.org/abs/2212.09720)
- [NVIDIA: Introducing NVFP4](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Red Hat: Accelerating large language models with NVFP4 quantization (2026)](https://developers.redhat.com/articles/2026/02/04/accelerating-large-language-models-nvfp4-quantization)

Interactive page: https://howaiworks.dev/quantization/
