Quantization-Aware Training
Quantization-aware training is a technique that accounts for low numeric precision during training to limit accuracy loss.
Quantization-aware training, or QAT, is a technique that builds quantization, the practice of storing a model's numbers at lower precision, into the training process. By simulating low-precision computation during training, it helps the model keep more of its accuracy after it is quantized.
Compared with post-training quantization, which compresses a model after training is finished, QAT usually preserves accuracy better but requires extra training. It is commonly used to run models on phones and small devices with limited memory.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.