quantization
Quantization is a technique that stores an AI model's weights as lower-bit numbers to reduce the model's size and memory use.
Quantization typically converts weights stored as 16-bit or 32-bit floating-point numbers into lower precision such as 8, 4 or 2 bits. It lets large models run with less memory and can make them faster, but output quality may drop as the bit count falls. It is widely used for local LLMs that run large language models on personal computers.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Q4 means 4-bit quantization: it stores weights at lower precision to reduce the file and memory footprint.