quantization
Shrinking a model by storing its numbers with less precision, so it runs faster and cheaper.
A model's weights are usually stored as very precise numbers. Quantization rounds them to less precise ones that take less memory. The model gets smaller and faster, and can run on cheaper hardware. It usually loses a little accuracy, so teams test it on their own tasks before switching.
Think of it as
Like saving a photo at lower quality to send it by message. It loads much faster and looks almost the same, unless you zoom right in.
Example
A model too big for one GPU is quantized so it fits. The team checks it on a couple of hundred of their real cases, finds the answers barely change, and moves to a smaller, cheaper machine.