multi-token prediction
Multi-token prediction is a technique in which a language model predicts several upcoming tokens at once instead of only the next one.
1 article
Last mentioned Often shortened to MTP, multi-token prediction is used during training to improve learning efficiency by predicting several tokens together, and during inference to draft tokens ahead of time so speculative decoding can speed up generation.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.