Quantization
Quantization stores a model's parameters with fewer digits of precision, making it much smaller and faster with only a small loss of quality. It is how large open models are squeezed onto laptops and phones.
In one line, for a 12-year-old
Quantization is shrinking an AI by rounding its numbers, like saving a photo at a smaller size.
An example
A "4-bit" version of an open model takes roughly a quarter of the memory of the original.
Why it matters to people
Smaller, local models help keep sensitive information on your own device instead of in someone else's cloud.