prefill

Prefill is the stage in which a large language model processes the entire input prompt at once before it starts generating a response.

1 article
Last mentioned

During prefill, the model computes all prompt tokens in parallel and builds the KV cache used for later generation. Unlike the following decode stage, which produces tokens one at a time, prefill is compute-heavy and depends largely on processing power. Its speed is usually expressed in tokens processed per second.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Uncached prefill 235.1 tokens per second at 512 tokens, 295.1 at 8k, 271.2 at 32k


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.