Floating point numbers are precise but heavy. Quantization reduces the numerical precision of model weights and activations, typically from 32-bit floats to 8-bit or 4-bit integers. The model shrinks, runs faster, and uses less power.
The trade-off is accuracy. Aggressive quantization degrades output quality, especially on tasks that need fine numerical distinctions. Post-training quantization is cheap but lossy. Quantization-aware training simulates low precision during training, recovering much of the lost accuracy.
Common quantization levels
- 16-bit floating point, minor compression
- 8-bit integer, widely used for inference
- 4-bit integer, aggressive but feasible with care
- Binary and ternary, experimental
Hardware support varies. Some accelerators handle 8-bit natively and deliver large speedups. Others fall back to slower paths. Choosing a quantization scheme depends as much on the deployment target as on the model.
Comments
No comments yet. Be the first to share a thought.
Leave a comment