prefill
Prefill is the stage in which a large language model processes the entire input prompt at once before it starts generating a response.
During prefill, the model computes all prompt tokens in parallel and builds the KV cache used for later generation. Unlike the following decode stage, which produces tokens one at a time, prefill is compute-heavy and depends largely on processing power. Its speed is usually expressed in tokens processed per second.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.