Big models win benchmarks. Small models ship. Model compression reduces size and compute while preserving most of the original accuracy, making deployment feasible on phones, embedded devices, and cost-sensitive servers.
Several techniques combine. Pruning removes weights that contribute little. Quantization reduces numerical precision. Distillation trains a smaller student. Low-rank factorization decomposes weight matrices. In practice, engineers stack these methods.
Common compression methods
- Pruning unimportant weights or neurons
- Quantization to 8-bit or lower
- Knowledge distillation to smaller students
- Low-rank matrix decomposition
- Architecture search for efficient designs
Compression has limits. Accuracy degrades, and aggressive methods can break behaviour in ways that only appear on rare inputs. Benchmarking on representative data before and after compression is essential.
Comments
No comments yet. Be the first to share a thought.
Leave a comment