Quantization

Inference & Serving Advanced

Quantization reduces the numeric precision used to store a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, shrinking memory use and often speeding up inference, at some cost to accuracy.

In simple terms

A model's weights are just numbers. Storing each one with fewer bits makes the whole model smaller in memory and faster to move around, the same way a lower-resolution image file is smaller and loads faster, at the cost of some lost detail.

Why it matters

It's one of the few techniques that can meaningfully shrink both memory footprint and inference cost for an existing, already-trained model, without retraining it from scratch.

How it works

Weights stored as 16-bit floating point numbers are converted to lower-precision formats, 8-bit or 4-bit integers, using a scale factor that maps the reduced range back to something close to the original values. Post-training quantization does this after training is complete; quantization-aware training simulates the reduced precision during training itself so the model adapts to it, usually preserving more accuracy.

Where it fits

Trained model (FP16 weights) → Quantization → Compressed weights (INT8 / INT4) → Smaller memory footprint → Faster inference

Production impact

Lower precision means less GPU memory used per model copy (more concurrent requests fit) and often faster computation, but too aggressive a quantization level measurably degrades output quality, a real accuracy-versus-cost trade-off, not a free win.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →